← Leaderboard
DeepSeek · Open weights · tested 8 Oct 2026
DeepSeek V4.1 Flash Max effort
Overall rank#4 of 11
Solves57/10050 first attempt · +7 retry
Rungs proven287/400254 first attempt · +33 retry
Cost to run$12.16$0.21 per solve
Tokens used813.9M96.2% from cache · 11.5M output
Time per lab18.5 minaverage of 149 attempts
Where it places
Each line is one metric across all 11 configurations. Right is always better; this model is the large dot.
Solveslabs solved, of 100
worsebetter
57#4 of 11
Rungs provenrungs proven, of 400
worsebetter
287#4 of 11
First-attempt solvessolved on the first attempt
worsebetter
50#4 of 11
Cost per solveUSD per solved lab
worsebetter
$0.21#2 of 7
By vulnerability class
How DeepSeek V4.1 Flash (Max) did on each type of vulnerability. Shown as counts, because some classes have only a few labs.
- Authentication / authorisation bypass17 of 28 solved87 of 112 rungs
- Information disclosure9 of 12 solved38 of 48 rungs
- IDOR5 of 11 solved32 of 44 rungs
- SSRF0 of 10 solved12 of 40 rungs
- Business logic6 of 10 solved32 of 40 rungs
- Key or secret exposure8 of 8 solved32 of 32 rungs
- XSS4 of 6 solved17 of 24 rungs
- Token theft5 of 5 solved20 of 20 rungs
Fewer than 5 labs each: one lab changes these a lot, so read them with care.
- Supply chain1 of 4 solved7 of 16 rungs
- Remote code execution0 of 3 solved1 of 12 rungs
- SQL injection2 of 2 solved8 of 8 rungs
- File inclusion0 of 1 solved1 of 4 rungs
Provider performance
Measured per single agent against the provider endpoint used for the run.
- Response time
- 10.35 s median · 26.79 s p95
- Output speed
- 80.4 tok/s
- API success rate
- 98.39%
- Input context
- 56.7K median · 157.1K p95
About this result
Every model runs in the same neutral harness, so results compare the models themselves.
- Harness
- BB Arena neutral harness
- Safety refusals
- None