Announcement

Code Index is live!

The first published revision of Revelo's Code Index is now available.

We introduce the first revision of Code Index: a live and evolving benchmark composed of novel tasks.

Our goal is to gather some of the most recognized coding benchmarks and provide new contamination-free results, so we can demonstrate how each model performs under different scenarios. This initial version includes tasks modeled on SWE-Bench-Pro Hard, SWE-Lancer, Terminal Bench 2.1, and Frontier-Bench.

To better understand the details of the methodology we applied, see Methodology.

Code Index leaderboard at revision 1, showing 13 model and agent pairings ranked by index score
Code Index leaderboard, revision 1. See the live board.

Main results

  • Claude Opus 5 is the overall leader. It showed great results especially on Terminal Bench 2.1 and Frontier-Bench tasks, with an outstanding increase in performance compared to the previous Claude Opus 4.8 (13.3% to 31.7% on Frontier-Bench).
  • GPT 5.6 Sol and Claude Fable 5 sit close to each other. They present very similar overall scores, trading positions across task sets. However, the average time and token costs favor GPT 5.6 Sol, which is the more efficient of the two.
  • Claude Opus 4.8, Grok 4.5, and Kimi K3 come in the next tier. Grok 4.5 had an outstanding performance on the SWE-Lancer benchmark, taking second place behind Claude Fable 5. Kimi K3 beats Claude Opus 4.8 on both the Terminal Bench 2.1 and Frontier-Bench task sets, despite costing much less.
  • The next tier brings GLM 5.2 and Gemini 3.6 Flash. In efficiency, they switch places: Gemini 3.6 Flash takes less time, but GLM 5.2 is cheaper for similar pass rates.
  • The remaining models compose the final tier, with Hy3 worth highlighting. It ties with DeepSeek V4 Pro for the best score in this group while being the cheapest of all measured models.
Efficiency charts plotting pass rate against mean agent time and mean cost per task
Pass rate against mean agent time and mean cost per task. Explore the efficiency charts.

Tasks with very high complexity

  • Most tasks fall on the hard side. Difficulty was measured empirically, based on the trials of all models for each task, and most of them presented a pass rate below 33%.
  • Eight tasks from the Frontier-Bench task set were not solved by any model. Only one task proved trivial, with every model succeeding on it.
  • Check the trial outcomes matrix to see each task’s result.
Difficulty distribution chart plotting the empirical difficulty of each task per task set.
Empirical difficulty measured for all 80 tasks based on the pass rates of the 13 measured models. Explore the difficulty distribution chart.

← Back to Announcements

Talk to us

Revelo works with model labs on:

  • Training datasets for code generation
  • Evaluations of new and unreleased models
  • Leaderboard analysis, including failure modes and agent trajectories

Talk to us

Tell us whether you need a dataset, an evaluation, or help analyzing model results.

Loading form…