Benchmarks
How we built our tasks
Code Index currently scores 80 private tasks across 4 benchmarks. Every single task is novel and contamination-free, designed specifically for this index, following rigorously the rubrics and format of each of the selected coding benchmarks. This dataset will continuously expand and update as new benchmarks are included and new, more challenging tasks are created.
Following an existing format makes results legible to anyone familiar with the original benchmark. Keeping the tasks private reduces the chance that a model encountered them during training. Each benchmark page describes their domain and difficulty distribution and presents an example task.
Composition
Every scored benchmark contributes equally to the index, regardless of task count.
| Benchmark | Modeled on | Tasks | Domains | Challenging | Hard | Medium | Easy | Trivial | In index |
|---|---|---|---|---|---|---|---|---|---|
| SWE-Bench-Pro Hard | Scale AI's SWE-Bench Pro | 20 | 1 | 0 | 7 | 9 | 3 | 1 | Yes |
| SWE-Lancer | OpenAI's SWE-Lancer | 20 | 1 | 0 | 10 | 9 | 1 | 0 | Yes |
| Terminal Bench 2.1 | Terminal Bench, by Stanford and the Laude Institute | 20 | 11 | 0 | 8 | 12 | 0 | 0 | Yes |
| Frontier-Bench | Frontier-Bench, by Stanford and the Laude Institute | 20 | 5 | 8 | 10 | 2 | 0 | 0 | Yes |
| Full-Stack WebApp | — | — | — | — | — | — | — | — | WIP |
| Infra as Code | — | — | — | — | — | — | — | — | WIP |
| UI Bench | — | — | — | — | — | — | — | — | WIP |
Difficulty is empirical and is read from the pass rate across measured models: challenging is 0%, hard is above 0% to 33%, medium is above 33% to 66%, easy is above 66% to below 100%, and trivial is 100%. Full methodology →
Talk to us
Revelo works with model labs on:
- Training datasets for code generation
- Evaluations of new and unreleased models
- Leaderboard analysis, including failure modes and agent trajectories