← Leaderboard

Cognition · Proprietary · tested 7 Oct 2026

SWE-2 Max effort

Overall rank#5 of 11
Solves52/10046 first attempt · +6 retry
Rungs proven258/400234 first attempt · +24 retry
Cost to runFree (promo)Free during Cognition's promotion. No list price yet, so it's left out of cost comparisons.
Tokens used618.2M96.5% from cache · 6.1M output
Time per lab18.5 minaverage of 149 attempts

Where it places

Each line is one metric across all 11 configurations. Right is always better; this model is the large dot.

Solveslabs solved, of 100
worsebetter
52#5 of 11
Rungs provenrungs proven, of 400
worsebetter
258#6 of 11
First-attempt solvessolved on the first attempt
worsebetter
46#5 of 11
Cost per solveUSD per solved lab
worsebetter
Free (promo)not ranked

By vulnerability class

How SWE-2 (Max) did on each type of vulnerability. Shown as counts, because some classes have only a few labs.

  • Authentication / authorisation bypass16 of 28 solved80 of 112 rungs
  • Information disclosure8 of 12 solved35 of 48 rungs
  • IDOR5 of 11 solved26 of 44 rungs
  • SSRF1 of 10 solved12 of 40 rungs
  • Business logic6 of 10 solved32 of 40 rungs
  • Key or secret exposure8 of 8 solved32 of 32 rungs
  • XSS2 of 6 solved11 of 24 rungs
  • Token theft5 of 5 solved20 of 20 rungs

Fewer than 5 labs each: one lab changes these a lot, so read them with care.

  • Supply chain0 of 4 solved4 of 16 rungs
  • Remote code execution0 of 3 solved1 of 12 rungs
  • SQL injection1 of 2 solved4 of 8 rungs
  • File inclusion0 of 1 solved1 of 4 rungs

Provider performance

Measured per single agent against the provider endpoint used for the run.

Response time
9.07 s median · 24.62 s p95
Output speed
42.5 tok/s
API success rate
98.49%
Input context
40.2K median · 123.2K p95

About this result

Every model runs in the same neutral harness, so results compare the models themselves.

Harness
BB Arena neutral harness
Safety refusals
None

See SWE-2 (Max) highlighted on every chart →