Announcement
Code Index is live!
The first published revision of Revelo's Code Index is now available.
We introduce the first revision of Code Index: a live and evolving benchmark composed of novel tasks.
Our goal is to gather some of the most recognized coding benchmarks and provide new contamination-free results, so we can demonstrate how each model performs under different scenarios. This initial version includes tasks modeled on SWE-Bench-Pro Hard, SWE-Lancer, Terminal Bench 2.1, and Frontier-Bench.
To better understand the details of the methodology we applied, see Methodology.
Main results
- Claude Opus 5 is the overall leader. It showed great results especially on Terminal Bench 2.1 and Frontier-Bench tasks, with an outstanding increase in performance compared to the previous Claude Opus 4.8 (13.3% to 31.7% on Frontier-Bench).
- GPT 5.6 Sol and Claude Fable 5 sit close to each other. They present very similar overall scores, trading positions across task sets. However, the average time and token costs favor GPT 5.6 Sol, which is the more efficient of the two.
- Claude Opus 4.8, Grok 4.5, and Kimi K3 come in the next tier. Grok 4.5 had an outstanding performance on the SWE-Lancer benchmark, taking second place behind Claude Fable 5. Kimi K3 beats Claude Opus 4.8 on both the Terminal Bench 2.1 and Frontier-Bench task sets, despite costing much less.
- The next tier brings GLM 5.2 and Gemini 3.6 Flash. In efficiency, they switch places: Gemini 3.6 Flash takes less time, but GLM 5.2 is cheaper for similar pass rates.
- The remaining models compose the final tier, with Hy3 worth highlighting. It ties with DeepSeek V4 Pro for the best score in this group while being the cheapest of all measured models.
Tasks with very high complexity
- Most tasks fall on the hard side. Difficulty was measured empirically, based on the trials of all models for each task, and most of them presented a pass rate below 33%.
- Eight tasks from the Frontier-Bench task set were not solved by any model. Only one task proved trivial, with every model succeeding on it.
- Check the trial outcomes matrix to see each task’s result.
Talk to us
Revelo works with model labs on:
- Training datasets for code generation
- Evaluations of new and unreleased models
- Leaderboard analysis, including failure modes and agent trajectories