For deeper analysis and transparency into the results, this page covers the metrics for each dataset and the results per trial for all 4 benchmarks.

Code Index per benchmark

Here we show not only the Index score per benchmark, but also the consumed tokens and the total cost of each model.

1
gpt-5.6-sol
codex
176.0M 985.5K
$142.9
86.7
2
claude-opus-5
claude-code
308.2M 2.3M
$259.3
83.3
3
claude-fable-5
claude-code
188.9M 1.9M
$347.9
71.7
4
grok-4.5
terminus-2
115.4M 2.2M
$59.2
60.0
5
claude-opus-4.8
claude-code
269.1M 1.9M
$214.2
58.3
6
kimi-k3
terminus-2
52.2M 1.4M
$45.0
50.0
7
glm-5.2
terminus-2
122.3M 2.5M
$70.2
41.7
8
gemini-3.6-flash
terminus-2
712.6M 4.4M
$236.3
40.0
9
hy3
terminus-2
193.3M 4.0M
$10.4
33.3
10
nemotron-3-ultra-550b-a55b
terminus-2
186.8M 2.7M
$58.2
28.3
11
mimo-v2.5-pro
terminus-2
141.2M 2.9M
$17.7
26.7
12
minimax-m3
terminus-2
472.3M 2.6M
$38.8
26.7
13
deepseek-v4-pro
terminus-2
178.9M 4.4M
$47.1
25.0

Trial outcomes

Below you can find the outcome of each trial. Each model ran three trials per task, which resulted in either PASS, FAIL, or ERROR. Each is described below, along with the error description, task domain, and task empirical difficulty, which can be seen by hovering each result dot.

Talk to us

Revelo works with model labs on:

  • Training datasets for code generation
  • Evaluations of new and unreleased models
  • Leaderboard analysis, including failure modes and agent trajectories

Talk to us

Tell us whether you need a dataset, an evaluation, or help analyzing model results.

Loading form…