← Leaderboard

OpenAI · Proprietary · tested 7 Oct 2026

GPT-5.6 Sol High effort

Overall rank#8 of 11
Solves46/10039 first attempt · +7 retry
Rungs proven250/400220 first attempt · +30 retry
Cost to run$95.97$2.09 per solve
Tokens used61.2M79.5% from cache · 1.6M output
Time per lab6.2 minaverage of 139 attempts

Where it places

Each line is one metric across all 11 configurations. Right is always better; this model is the large dot.

Solveslabs solved, of 100
worsebetter
46#8 of 11
Rungs provenrungs proven, of 400
worsebetter
250#7 of 11
First-attempt solvessolved on the first attempt
worsebetter
39#7 of 11
Cost per solveUSD per solved lab
worsebetter
$2.09#3 of 7

By vulnerability class

How GPT-5.6 Sol (High) did on each type of vulnerability. Shown as counts, because some classes have only a few labs.

  • Authentication / authorisation bypass17 of 28 solved82 of 112 rungs
  • Information disclosure8 of 12 solved35 of 48 rungs
  • IDOR3 of 11 solved21 of 44 rungs
  • SSRF0 of 10 solved9 of 40 rungs
  • Business logic6 of 10 solved31 of 40 rungs
  • Key or secret exposure7 of 8 solved30 of 32 rungs
  • XSS4 of 6 solved16 of 24 rungs
  • Token theft1 of 5 solved14 of 20 rungs

Fewer than 5 labs each: one lab changes these a lot, so read them with care.

  • Supply chain0 of 4 solved8 of 16 rungs
  • Remote code execution0 of 3 solved1 of 12 rungs
  • SQL injection0 of 2 solved3 of 8 rungs
  • File inclusion0 of 1 solved0 of 4 rungs

Provider performance

Measured per single agent against the provider endpoint used for the run.

Response time
11.42 s median · 40.87 s p95
Output speed
35.1 tok/s
API success rate
91.31%
Input context
15.5K median · 46.6K p95

About this result

Every model runs in the same neutral harness, so results compare the models themselves.

Harness
BB Arena neutral harness
Safety refusals
None

See GPT-5.6 Sol (High) highlighted on every chart →