Code Index
Code Index is a live and evolving coding-benchmarks leaderboard composed of novel, contamination-free collections. It currently combines 4 private task sets from traditional benchmarks: SWE-Bench-Pro Hard, SWE-Lancer, Terminal Bench 2.1, Frontier-Bench. More task sets soon to come, including proprietary benchmarks.
Code Index leaderboard
The Index is the mean of each benchmark's pass rates, each with equal weight. Read the methodology.
Per benchmark
Our results show that models can have different standings for different benchmarks. See the detailed results per benchmark.
Three trials ran for each task and the mean of all 'n' tasks composes the per-benchmark index. Read the methodology.
Efficiency tradeoffs
The frontier is not only the highest-scoring models, but those that can achieve best results more efficiently. The following charts present the mean agent time and token cost for each model, and which of them compose the frontier line. Better results appear toward the upper left.
Score vs. steps
Tha number of steps needed to reach a given Index score shows which models make the most of their earlier agent steps, which may be crucial under time/cost-limited scenarios.
Accessible summary
| Model | Score at ≤ ceil(P95) steps | ceil(P95) steps |
|---|---|---|
| claude-opus-5 | 54.2% | 105 |
| gpt-5.6-sol | 48.8% | 66 |
| claude-fable-5 | 47.5% | 58 |
| claude-opus-4.8 | 39.6% | 91 |
| kimi-k3 | 37.9% | 66 |
| grok-4.5 | 33.8% | 70 |
| glm-5.2 | 29.2% | 94 |
| gemini-3.6-flash | 26.7% | 185 |
| deepseek-v4-pro | 18.3% | 78 |
| hy3 | 18.3% | 75 |
| minimax-m3 | 16.7% | 150 |
| mimo-v2.5-pro | 10.8% | 64 |
| nemotron-3-ultra-550b-a55b | 8.8% | 128 |
Difficulty distribution
Difficulty was categorized according to the results of each task across all measured models. Full methodology →
Talk to us
Revelo works with model labs on:
- Training datasets for code generation
- Evaluations of new and unreleased models
- Leaderboard analysis, including failure modes and agent trajectories