Code Index currently scores 80 private tasks across 4 benchmarks. Every single task is novel and contamination-free, designed specifically for this index, following rigorously the rubrics and format of each of the selected coding benchmarks. This dataset will continuously expand and update as new benchmarks are included and new, more challenging tasks are created.

Following an existing format makes results legible to anyone familiar with the original benchmark. Keeping the tasks private reduces the chance that a model encountered them during training. Each benchmark page describes their domain and difficulty distribution and presents an example task.

Composition

Every scored benchmark contributes equally to the index, regardless of task count.

Benchmark Modeled on Tasks Domains Challenging Hard Medium Easy Trivial In index
SWE-Bench-Pro Hard Scale AI's SWE-Bench Pro 20 1 0 7 9 3 1 Yes
SWE-Lancer OpenAI's SWE-Lancer 20 1 0 10 9 1 0 Yes
Terminal Bench 2.1 Terminal Bench, by Stanford and the Laude Institute 20 11 0 8 12 0 0 Yes
Frontier-Bench Frontier-Bench, by Stanford and the Laude Institute 20 5 8 10 2 0 0 Yes
Full-Stack WebApp WIP
Infra as Code WIP
UI Bench WIP

Difficulty is empirical and is read from the pass rate across measured models: challenging is 0%, hard is above 0% to 33%, medium is above 33% to 66%, easy is above 66% to below 100%, and trivial is 100%. Full methodology →

Talk to us

Revelo works with model labs on:

  • Training datasets for code generation
  • Evaluations of new and unreleased models
  • Leaderboard analysis, including failure modes and agent trajectories

Talk to us

Tell us whether you need a dataset, an evaluation, or help analyzing model results.

Loading form…