bbarena v1.0 · updated 8 Oct 2026

Which AI models can find and prove real vulnerabilities?

BB Arena runs each AI model as an autonomous agent, black-box, against 100 web and API labs built from original security research. No source code, no hints, and private labs that models can't have trained on. Each lab is a four-rung ladder, from finding the weak spot to completing the exploit, and only counts as solved when the agent proves all four.

Grok 4.6 (High) leads with 62 of 100 labs solved, ahead of Claude Opus 5 (High) on 60.

All models
    • 60
    • 48
    • 46
    • 35
    • 52
    • 57
    • 46
    • 0
    • 49
    • 62
    • 59
Colour

Score

How many of the 100 black-box labs each model fully exploited and proved.

labs solved, of 100·Higher is better
  • Anthropic
  • Cognition
  • DeepSeek
  • OpenAI
  • Xiaomi
  • xAI
  • First attempt
  • Added on retry
What this measures

Labs where the agent proved all four rungs. A rung counts as proven only when the hidden verifier saw the action and the agent's own trace contains the exact proof for that lab instance. Final result after the first attempt and the retry.

Safety refusals

labs with refusals, of 100 · Lower is better

Labs where the provider's cyber safeguards stopped at least one attempt, by refusing or by answering with a different model. Stopped attempts score zero: we never count a fallback model's work. Models with no refusals aren't listed.

How refusals are scored

Leaderboard

Every configuration tested: one model at one effort level, all in the same neutral harness.

Model
1Grok 4.6 HighxAI · Proprietary623020$146$2.35145.2M17.2 min53.5 tok/s97.87%
2Claude Opus 5 HighAnthropic · Proprietary602930$463$7.72312.5M17.1 min65.7 tok/s99.08%
3Grok 4.7 HighxAI · Proprietary592950$290$4.92453.3M8.7 min85.3 tok/s99.88%
4DeepSeek V4.1 Flash MaxDeepSeek · Open weights572870$12.16$0.21813.9M18.5 min80.4 tok/s98.39%
5SWE-2 MaxCognition · Proprietary522580Free (promo)n/a618.2M18.5 min42.5 tok/s98.49%
6MiMo V2.6 Flash HighXiaomi · Open weights492600$3.43$0.07181.9M20.4 min31.2 tok/s94.52%
7Claude Mythos 5.1 HighSafety refusalsAnthropic · Proprietary4822452n/an/an/an/a62.9 tok/s99.56%
8GPT-5.6 Sol HighOpenAI · Proprietary462500$95.97$2.0961.2M6.2 min35.1 tok/s91.31%
9Claude Sonnet 5 HighAnthropic · Proprietary462430$258$5.60587.3M18.2 min54.4 tok/s99.39%
10Claude Opus 5.5 HighSafety refusalsAnthropic · Proprietary3518651n/an/an/an/a73.7 tok/s99.97%
11GPT-6 Astra HighOpenAI · Proprietary0170$36.34n/a3.9M1.4 min25.2 tok/s91.34%

Ranked by solves (of 100); rungs proven break ties. Bar: first attempt, lighter segment = added on retry. Click a column to sort, a model for its full profile. Refusals: labs where the provider's cyber safeguards stopped an attempt. Stopped attempts score zero and we never count a fallback model's work; with many refusals, cost, tokens and time are left out of comparisons (why).

By vulnerability class

Labs solved per type of vulnerability, to compare models where it matters to you. Each model's page has the full detail.

ModelAuth bypass28 labsInfo disclosure12 labsIDOR11 labsSSRF10 labsBusiness logic10 labsKey exposure8 labsXSS6 labsToken theft5 labsSupply chain4 labsRCE3 labsSQL injection2 labsFile inclusion1 lab
Grok 4.6 High19/2810/127/110/106/108/85/63/53/40/31/20/1
Claude Opus 5 High17/2810/123/112/105/108/84/65/54/40/32/20/1
Grok 4.7 High18/2810/126/111/106/108/84/65/50/40/31/20/1
DeepSeek V4.1 Flash Max17/289/125/110/106/108/84/65/51/40/32/20/1
SWE-2 Max16/288/125/111/106/108/82/65/50/40/31/20/1
MiMo V2.6 Flash High16/288/122/110/106/108/84/63/51/40/31/20/1
Claude Mythos 5.1 High16/288/126/110/108/106/84/60/50/40/30/20/1
GPT-5.6 Sol High17/288/123/110/106/107/84/61/50/40/30/20/1
Claude Sonnet 5 High18/288/123/110/106/103/84/64/50/40/30/20/1
Claude Opus 5.5 High13/284/122/110/106/108/81/60/50/40/31/20/1
GPT-6 Astra High0/280/120/110/100/100/80/60/50/40/30/20/1
0% solved100%Best in classFaded columns have fewer than 5 labs, so one lab changes them a lot.Safety refusals in some of that class's labs; stopped attempts score zero (why).

Value

Solves against what it took to get them.

Up and to the left is better
  • Anthropic
  • Cognition
  • DeepSeek
  • OpenAI
  • Xiaomi
  • xAI

Cost & tokens

What it cost to run the full benchmark, at list prices on the test date. Cost per solve accounts for results; cost to run on its own does not.

USD per solved lab·Lower is better
  • Anthropic
  • Cognition
  • DeepSeek
  • OpenAI
  • Xiaomi
  • xAI
What this measures

Cost to run divided by solves: what one fully proven finding cost.

Speed & reliability

Measured per single agent, so the numbers reflect the model and its provider, not how many agents we ran in parallel.

avg minutes per lab attempt·Lower is faster, not necessarily better
  • Anthropic
  • Cognition
  • DeepSeek
  • OpenAI
  • Xiaomi
  • xAI
What this measures

Average wall-clock time one agent spent on one lab attempt. Each attempt is capped at 25 minutes. A quick time can also mean the agent gave up early, so read it alongside solves.

Why BB Arena is different

Most cyber benchmarks hand the model source code or reuse public challenges. Real attackers get neither, and public tasks can leak into training data.

What the model gets
Source code to read (white-box), or puzzle-style capture-the-flag challenges.
A target, a scope and a shell. Black-box, like a real external security test.
The agent harness
Each model in its vendor's own agent tool, or a harness tuned for one model family.
One neutral harness for every model: the same tool, instructions, context handling and limits.
Where the tasks come from
Public tasks, often with solutions and write-ups published online.
100 fully synthetic labs built from our own security research into real-world vulnerabilities.
Training-data contamination
Public tasks and answers can end up in training data, so scores can rise without real skill improving.
Labs, prompts and verifiers are never published, and proof values change on every run, so there are no answers to learn.
What counts as success
Often the model's final answer or a single flag.
Verified proof at every step: the exact proof value and a server-recorded action, rung by rung.
Want your model on BB Arena?AI labs, providers and sponsors: we evaluate public and upcoming models under the same protocol as every other model.
Get your model tested →

How a lab is scored

Each lab recreates a real-world vulnerability in a fully synthetic, isolated environment. The agent gets a target, a scope and a shell. Nothing else.

  1. 1Establish the surfaceFind the endpoint, feature or precondition the bug lives in.
  2. 2Cross the first boundaryBreak the first security control: another user's object, an internal host, a privileged route.
  3. 3Go deeperTurn the foothold into access to consequential data or state.
  4. 4Complete the chainFinish the full exploit. All four proven = one solve.

A rung counts as proven only when our hidden verifier saw the action and the agent's own trace contains that lab instance's unique proof. What the agent claims earns nothing. Every configuration gets up to two attempts at each lab; the retry only runs on labs missed the first time, and starts fresh. Full methodology →