bbarena v1.0 · updated 8 Oct 2026
Which AI models can find and prove real vulnerabilities?
BB Arena runs each AI model as an autonomous agent, black-box, against 100 web and API labs built from original security research. No source code, no hints, and private labs that models can't have trained on. Each lab is a four-rung ladder, from finding the weak spot to completing the exploit, and only counts as solved when the agent proves all four.
Grok 4.6 (High) leads with 62 of 100 labs solved, ahead of Claude Opus 5 (High) on 60.
All models
- 60
- 48
- 46
- 35
- 52
- 57
- 46
- 0
- 49
- 62
- 59
- No matching models
- Grok 4.6 High62 solves
- Claude Opus 5 High60 solves
- Grok 4.7 High59 solves
- DeepSeek V4.1 Flash Max57 solves
- SWE-2 Max52 solves
- MiMo V2.6 Flash High49 solves
- Claude Mythos 5.1 High48 solves
- GPT-5.6 Sol High46 solves
- Claude Sonnet 5 High46 solves
- Claude Opus 5.5 High35 solves
- GPT-6 Astra High0 solves
Score
How many of the 100 black-box labs each model fully exploited and proved.
- Anthropic
- Cognition
- DeepSeek
- OpenAI
- Xiaomi
- xAI
- First attempt
- Added on retry
What this measures
Labs where the agent proved all four rungs. A rung counts as proven only when the hidden verifier saw the action and the agent's own trace contains the exact proof for that lab instance. Final result after the first attempt and the retry.
Safety refusals
labs with refusals, of 100 · Lower is betterLabs where the provider's cyber safeguards stopped at least one attempt, by refusing or by answering with a different model. Stopped attempts score zero: we never count a fallback model's work. Models with no refusals aren't listed.
- Claude Mythos 5.1 High52 of 100 labsAnthropic refused or answered with Claude Opus 5 instead
- Claude Opus 5.5 High51 of 100 labsAnthropic answered with Claude Opus 4.8 instead
Leaderboard
Every configuration tested: one model at one effort level, all in the same neutral harness.
| Model | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Grok 4.6 HighxAI · Proprietary | 62 | 302 | 0 | $146 | $2.35 | 145.2M | 17.2 min | 53.5 tok/s | 97.87% |
| 2 | Claude Opus 5 HighAnthropic · Proprietary | 60 | 293 | 0 | $463 | $7.72 | 312.5M | 17.1 min | 65.7 tok/s | 99.08% |
| 3 | Grok 4.7 HighxAI · Proprietary | 59 | 295 | 0 | $290 | $4.92 | 453.3M | 8.7 min | 85.3 tok/s | 99.88% |
| 4 | DeepSeek V4.1 Flash MaxDeepSeek · Open weights | 57 | 287 | 0 | $12.16 | $0.21 | 813.9M | 18.5 min | 80.4 tok/s | 98.39% |
| 5 | SWE-2 MaxCognition · Proprietary | 52 | 258 | 0 | Free (promo) | n/a | 618.2M | 18.5 min | 42.5 tok/s | 98.49% |
| 6 | MiMo V2.6 Flash HighXiaomi · Open weights | 49 | 260 | 0 | $3.43 | $0.07 | 181.9M | 20.4 min | 31.2 tok/s | 94.52% |
| 7 | Claude Mythos 5.1 HighSafety refusalsAnthropic · Proprietary | 48 | 224 | 52 | n/a | n/a | n/a | n/a | 62.9 tok/s | 99.56% |
| 8 | GPT-5.6 Sol HighOpenAI · Proprietary | 46 | 250 | 0 | $95.97 | $2.09 | 61.2M | 6.2 min | 35.1 tok/s | 91.31% |
| 9 | Claude Sonnet 5 HighAnthropic · Proprietary | 46 | 243 | 0 | $258 | $5.60 | 587.3M | 18.2 min | 54.4 tok/s | 99.39% |
| 10 | Claude Opus 5.5 HighSafety refusalsAnthropic · Proprietary | 35 | 186 | 51 | n/a | n/a | n/a | n/a | 73.7 tok/s | 99.97% |
| 11 | GPT-6 Astra HighOpenAI · Proprietary | 0 | 17 | 0 | $36.34 | n/a | 3.9M | 1.4 min | 25.2 tok/s | 91.34% |
Ranked by solves (of 100); rungs proven break ties. Bar: first attempt, lighter segment = added on retry. Click a column to sort, a model for its full profile. Refusals: labs where the provider's cyber safeguards stopped an attempt. Stopped attempts score zero and we never count a fallback model's work; with many refusals, cost, tokens and time are left out of comparisons (why).
By vulnerability class
Labs solved per type of vulnerability, to compare models where it matters to you. Each model's page has the full detail.
| Model | Auth bypass28 labs | Info disclosure12 labs | IDOR11 labs | SSRF10 labs | Business logic10 labs | Key exposure8 labs | XSS6 labs | Token theft5 labs | Supply chain4 labs | RCE3 labs | SQL injection2 labs | File inclusion1 lab |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Grok 4.6 High | 19/28 | 10/12 | 7/11 | 0/10 | 6/10 | 8/8 | 5/6 | 3/5 | 3/4 | 0/3 | 1/2 | 0/1 |
| Claude Opus 5 High | 17/28 | 10/12 | 3/11 | 2/10 | 5/10 | 8/8 | 4/6 | 5/5 | 4/4 | 0/3 | 2/2 | 0/1 |
| Grok 4.7 High | 18/28 | 10/12 | 6/11 | 1/10 | 6/10 | 8/8 | 4/6 | 5/5 | 0/4 | 0/3 | 1/2 | 0/1 |
| DeepSeek V4.1 Flash Max | 17/28 | 9/12 | 5/11 | 0/10 | 6/10 | 8/8 | 4/6 | 5/5 | 1/4 | 0/3 | 2/2 | 0/1 |
| SWE-2 Max | 16/28 | 8/12 | 5/11 | 1/10 | 6/10 | 8/8 | 2/6 | 5/5 | 0/4 | 0/3 | 1/2 | 0/1 |
| MiMo V2.6 Flash High | 16/28 | 8/12 | 2/11 | 0/10 | 6/10 | 8/8 | 4/6 | 3/5 | 1/4 | 0/3 | 1/2 | 0/1 |
| Claude Mythos 5.1 High | 16/28 | 8/12 | 6/11 | 0/10 | 8/10 | 6/8 | 4/6 | 0/5 | 0/4 | 0/3 | 0/2 | 0/1 |
| GPT-5.6 Sol High | 17/28 | 8/12 | 3/11 | 0/10 | 6/10 | 7/8 | 4/6 | 1/5 | 0/4 | 0/3 | 0/2 | 0/1 |
| Claude Sonnet 5 High | 18/28 | 8/12 | 3/11 | 0/10 | 6/10 | 3/8 | 4/6 | 4/5 | 0/4 | 0/3 | 0/2 | 0/1 |
| Claude Opus 5.5 High | 13/28 | 4/12 | 2/11 | 0/10 | 6/10 | 8/8 | 1/6 | 0/5 | 0/4 | 0/3 | 1/2 | 0/1 |
| GPT-6 Astra High | 0/28 | 0/12 | 0/11 | 0/10 | 0/10 | 0/8 | 0/6 | 0/5 | 0/4 | 0/3 | 0/2 | 0/1 |
Value
Solves against what it took to get them.
- Anthropic
- Cognition
- DeepSeek
- OpenAI
- Xiaomi
- xAI
Models with heavy safety refusals aren't shown in this chart. See Safety refusals
Cost & tokens
What it cost to run the full benchmark, at list prices on the test date. Cost per solve accounts for results; cost to run on its own does not.
- Anthropic
- Cognition
- DeepSeek
- OpenAI
- Xiaomi
- xAI
Models with heavy safety refusals aren't shown in this chart. See Safety refusals
What this measures
Cost to run divided by solves: what one fully proven finding cost.
Speed & reliability
Measured per single agent, so the numbers reflect the model and its provider, not how many agents we ran in parallel.
- Anthropic
- Cognition
- DeepSeek
- OpenAI
- Xiaomi
- xAI
Models with heavy safety refusals aren't shown in this chart. See Safety refusals
What this measures
Average wall-clock time one agent spent on one lab attempt. Each attempt is capped at 25 minutes. A quick time can also mean the agent gave up early, so read it alongside solves.
Why BB Arena is different
Most cyber benchmarks hand the model source code or reuse public challenges. Real attackers get neither, and public tasks can leak into training data.
- What the model gets
- Source code to read (white-box), or puzzle-style capture-the-flag challenges.
- A target, a scope and a shell. Black-box, like a real external security test.
- The agent harness
- Each model in its vendor's own agent tool, or a harness tuned for one model family.
- One neutral harness for every model: the same tool, instructions, context handling and limits.
- Where the tasks come from
- Public tasks, often with solutions and write-ups published online.
- 100 fully synthetic labs built from our own security research into real-world vulnerabilities.
- Training-data contamination
- Public tasks and answers can end up in training data, so scores can rise without real skill improving.
- Labs, prompts and verifiers are never published, and proof values change on every run, so there are no answers to learn.
- What counts as success
- Often the model's final answer or a single flag.
- Verified proof at every step: the exact proof value and a server-recorded action, rung by rung.
How a lab is scored
Each lab recreates a real-world vulnerability in a fully synthetic, isolated environment. The agent gets a target, a scope and a shell. Nothing else.
- 1Establish the surfaceFind the endpoint, feature or precondition the bug lives in.
- 2Cross the first boundaryBreak the first security control: another user's object, an internal host, a privileged route.
- 3Go deeperTurn the foothold into access to consequential data or state.
- 4Complete the chainFinish the full exploit. All four proven = one solve.
A rung counts as proven only when our hidden verifier saw the action and the agent's own trace contains that lab instance's unique proof. What the agent claims earns nothing. Every configuration gets up to two attempts at each lab; the retry only runs on labs missed the first time, and starts fresh. Full methodology →