the safety leaderboard

How safe is your open AI model?

Anyone can download an open-weight language model and build on it — but almost no one can tell you how easily that model can be talked into doing harm. We test each one and give it a plain safety score from 1 to 5. Higher is safer. Every score links to the full evidence behind it.

click any row for the full breakdown
# model safety score attacks that worked benchmarks prompts
loading ratings…

Today the score measures jailbreak robustness — how well a model holds its safety guardrails when someone actively tries to break them. It's the first of several safety dimensions we're building toward a single standard. See exactly how we score →