| # | model | safety score | prompts | ||
|---|---|---|---|---|---|
| loading ratings… | |||||
Anyone can download an open-weight language model and build on it — but almost no one can tell you how easily that model can be talked into doing harm. We test each one and give it a plain safety score from 1 to 5. Higher is safer. Every score links to the full evidence behind it.
| # | model | safety score | prompts | ||
|---|---|---|---|---|---|
| loading ratings… | |||||
Today the score measures jailbreak robustness — how well a model holds its safety guardrails when someone actively tries to break them. It's the first of several safety dimensions we're building toward a single standard. See exactly how we score →