← Leaderboard
Cognition · Proprietary · tested 7 Oct 2026
SWE-2 Max effort
Overall rank#5 of 11
Solves52/10046 first attempt · +6 retry
Rungs proven258/400234 first attempt · +24 retry
Cost to runFree (promo)Free during Cognition's promotion. No list price yet, so it's left out of cost comparisons.
Tokens used618.2M96.5% from cache · 6.1M output
Time per lab18.5 minaverage of 149 attempts
Where it places
Each line is one metric across all 11 configurations. Right is always better; this model is the large dot.
Solveslabs solved, of 100
worsebetter
52#5 of 11
Rungs provenrungs proven, of 400
worsebetter
258#6 of 11
First-attempt solvessolved on the first attempt
worsebetter
46#5 of 11
Cost per solveUSD per solved lab
worsebetter
Free (promo)not ranked
By vulnerability class
How SWE-2 (Max) did on each type of vulnerability. Shown as counts, because some classes have only a few labs.
- Authentication / authorisation bypass16 of 28 solved80 of 112 rungs
- Information disclosure8 of 12 solved35 of 48 rungs
- IDOR5 of 11 solved26 of 44 rungs
- SSRF1 of 10 solved12 of 40 rungs
- Business logic6 of 10 solved32 of 40 rungs
- Key or secret exposure8 of 8 solved32 of 32 rungs
- XSS2 of 6 solved11 of 24 rungs
- Token theft5 of 5 solved20 of 20 rungs
Fewer than 5 labs each: one lab changes these a lot, so read them with care.
- Supply chain0 of 4 solved4 of 16 rungs
- Remote code execution0 of 3 solved1 of 12 rungs
- SQL injection1 of 2 solved4 of 8 rungs
- File inclusion0 of 1 solved1 of 4 rungs
Provider performance
Measured per single agent against the provider endpoint used for the run.
- Response time
- 9.07 s median · 24.62 s p95
- Output speed
- 42.5 tok/s
- API success rate
- 98.49%
- Input context
- 40.2K median · 123.2K p95
About this result
Every model runs in the same neutral harness, so results compare the models themselves.
- Harness
- BB Arena neutral harness
- Safety refusals
- None