Code Index is a live and evolving coding-benchmarks leaderboard composed of novel, contamination-free collections. It currently combines 4 private task sets from traditional benchmarks: SWE-Bench-Pro Hard, SWE-Lancer, Terminal Bench 2.1, Frontier-Bench. More task sets soon to come, including proprietary benchmarks.

Code Index leaderboard

1
claude-opus-5
claude-code
1.35B 22.3M
$1.5k
57.1
2
gpt-5.6-sol
codex
469.2M 6.1M
$492.7
51.2
3
claude-fable-5
claude-code
770.6M 15.6M
$1.9k
50.0
4
claude-opus-4.8
claude-code
987.5M 13.4M
$967.1
41.2
5
kimi-k3
terminus-2
283.2M 9.5M
$268.8
39.6
6
grok-4.5
terminus-2
1.64B 14.5M
$900.7
35.8
7
glm-5.2
terminus-2
727.8M 24.0M
$446.0
30.8
8
gemini-3.6-flash
terminus-2
3.73B 17.3M
$987.5
28.3
9
deepseek-v4-pro
terminus-2
2.13B 33.0M
$365.3
19.6
10
hy3
terminus-2
835.6M 25.5M
$60.6
19.6
11
minimax-m3
terminus-2
2.34B 16.6M
$182.6
17.9
12
mimo-v2.5-pro
terminus-2
642.7M 17.1M
$117.7
11.2
13
nemotron-3-ultra-550b-a55b
terminus-2
1.05B 13.5M
$331.0
9.2

The Index is the mean of each benchmark's pass rates, each with equal weight. Read the methodology.

Per benchmark

Our results show that models can have different standings for different benchmarks. See the detailed results per benchmark.

SWE-Bench-Pro Hard n=20
claude-opus-5
83.3
gpt-5.6-sol
86.7
claude-fable-5
71.7
claude-opus-4.8
58.3
kimi-k3
50.0
grok-4.5
60.0
glm-5.2
41.7
gemini-3.6-flash
40.0
deepseek-v4-pro
25.0
hy3
33.3
minimax-m3
26.7
mimo-v2.5-pro
26.7
nemotron-3-ultra-550b-a55b
28.3
SWE-Lancer n=20
claude-opus-5
45.0
gpt-5.6-sol
43.3
claude-fable-5
51.7
claude-opus-4.8
45.0
kimi-k3
43.3
grok-4.5
46.7
glm-5.2
40.0
gemini-3.6-flash
41.7
deepseek-v4-pro
26.7
hy3
16.7
minimax-m3
21.7
mimo-v2.5-pro
6.7
nemotron-3-ultra-550b-a55b
1.7
Terminal Bench 2.1 n=20
claude-opus-5
68.3
gpt-5.6-sol
58.3
claude-fable-5
51.7
claude-opus-4.8
48.3
kimi-k3
50.0
grok-4.5
35.0
glm-5.2
38.3
gemini-3.6-flash
23.3
deepseek-v4-pro
21.7
hy3
21.7
minimax-m3
20.0
mimo-v2.5-pro
8.3
nemotron-3-ultra-550b-a55b
5.0
Frontier-Bench n=20
claude-opus-5
31.7
gpt-5.6-sol
16.7
claude-fable-5
25.0
claude-opus-4.8
13.3
kimi-k3
15.0
grok-4.5
1.7
glm-5.2
3.3
gemini-3.6-flash
8.3
deepseek-v4-pro
5.0
hy3
6.7
minimax-m3
3.3
mimo-v2.5-pro
3.3
nemotron-3-ultra-550b-a55b
1.7

Three trials ran for each task and the mean of all 'n' tasks composes the per-benchmark index. Read the methodology.

Efficiency tradeoffs

The frontier is not only the highest-scoring models, but those that can achieve best results more efficiently. The following charts present the mean agent time and token cost for each model, and which of them compose the frontier line. Better results appear toward the upper left.

Score vs. steps

Tha number of steps needed to reach a given Index score shows which models make the most of their earlier agent steps, which may be crucial under time/cost-limited scenarios.

Accessible summary

Score versus steps summary
Model Score at ≤ ceil(P95) steps ceil(P95) steps
claude-opus-5 54.2% 105
gpt-5.6-sol 48.8% 66
claude-fable-5 47.5% 58
claude-opus-4.8 39.6% 91
kimi-k3 37.9% 66
grok-4.5 33.8% 70
glm-5.2 29.2% 94
gemini-3.6-flash 26.7% 185
deepseek-v4-pro 18.3% 78
hy3 18.3% 75
minimax-m3 16.7% 150
mimo-v2.5-pro 10.8% 64
nemotron-3-ultra-550b-a55b 8.8% 128

Difficulty distribution

Difficulty was categorized according to the results of each task across all measured models. Full methodology →

Talk to us

Revelo works with model labs on:

  • Training datasets for code generation
  • Evaluations of new and unreleased models
  • Leaderboard analysis, including failure modes and agent trajectories

Talk to us

Tell us whether you need a dataset, an evaluation, or help analyzing model results.

Loading form…