Anthropic · Proprietary · tested 8 Oct 2026
Claude Mythos 5.1 High effort
Safety refusals affected this result
Anthropic's cyber safeguards stopped attempts in 52 of 100 labs, refusing or answering with Claude Opus 5 instead. Stopped attempts score zero. BB Arena never counts a fallback model's work: the result is this model's own. BB Arena is enrolled in the cyber programs the frontier AI labs offer to verified security researchers, including Anthropic's, and these refusals happened anyway.
Its cost, tokens and time don't cover the full benchmark, so they're left out of comparisons. How refusals are scored
Where it places
Each line is one metric across all 11 configurations. Right is always better; this model is the large dot.
By vulnerability class
How Claude Mythos 5.1 (High) did on each type of vulnerability. Shown as counts, because some classes have only a few labs.
- Authentication / authorisation bypassSafety refusals in 14 of 28 labs16 of 28 solved79 of 112 rungs
- Information disclosureSafety refusals in 5 of 12 labs8 of 12 solved34 of 48 rungs
- IDORSafety refusals in 6 of 11 labs6 of 11 solved31 of 44 rungs
- SSRFSafety refusals in 9 of 10 labs0 of 10 solved3 of 40 rungs
- Business logicSafety refusals in 2 of 10 labs8 of 10 solved35 of 40 rungs
- Key or secret exposureSafety refusals in 3 of 8 labs6 of 8 solved24 of 32 rungs
- XSSSafety refusals in 1 of 6 labs4 of 6 solved17 of 24 rungs
- Token theftSafety refusals in 5 of 5 labs0 of 5 solved0 of 20 rungs
Fewer than 5 labs each: one lab changes these a lot, so read them with care.
- Supply chainSafety refusals in 4 of 4 labs0 of 4 solved0 of 16 rungs
- Remote code execution0 of 3 solved1 of 12 rungs
- SQL injectionSafety refusals in 2 of 2 labs0 of 2 solved0 of 8 rungs
- File inclusionSafety refusals in 1 of 1 lab0 of 1 solved0 of 4 rungs
Provider performance
Measured per single agent against the provider endpoint used for the run.
- Response time
- 9.03 s median · 32.17 s p95
- Output speed
- 62.9 tok/s
- API success rate
- 99.56%
- Input context
- 19.7K median · 68.8K p95
About this result
Every model runs in the same neutral harness, so results compare the models themselves.
- Harness
- BB Arena neutral harness
- Safety refusals
- 52 of 100 labs