# Revelo Research > Independent research on AI systems and the work they produce. Revelo Research > publishes the Code Index, an execution-graded index of model-agent performance > on private software engineering benchmarks. Revelo builds code datasets and evaluations, and works with model labs on training data for code generation, evaluations of new and unreleased models, and leaderboard analysis. ## Code Index - [Code Index](https://research.revelo.com/code-index/): The live leaderboard. Ranks model-agent pairs by index score across 4 private, contamination-free benchmarks. Reports pass rate, cost, token usage, and mean agent time. - [Results by benchmark](https://research.revelo.com/code-index/results/): Per-benchmark boards, efficiency tradeoffs, and difficulty distribution. - [Methodology](https://research.revelo.com/code-index/methodology/): How the index is computed, graded, and authored. - [Benchmarks](https://research.revelo.com/code-index/benchmarks/): What the task sets contain and which public benchmarks they follow. ## Benchmarks - [SWE-Bench-Pro Hard](https://research.revelo.com/code-index/benchmarks/swe-bench-pro-hard/): Private repository bug-fixing tasks with held-out tests for the requested change and existing behavior.- [SWE-Lancer](https://research.revelo.com/code-index/benchmarks/swe-lancer/): Private freelance engineering tasks with original client specifications, executable tests, and closing prices.- [Terminal Bench 2.1](https://research.revelo.com/code-index/benchmarks/terminal-bench-2-1/): Private multi-step command-line tasks completed in isolated containers and graded from their final state.- [Terminal Bench 3.0](https://research.revelo.com/code-index/benchmarks/terminal-bench-3-0/): Private software engineering tasks evaluated in isolated command-line environments. ## Updates - [Announcements](https://research.revelo.com/code-index/announcements/): Releases, benchmark additions, and index updates. - [Qwen 3.8 Max and DeepSeek v4 Flash 0731 added!](https://research.revelo.com/code-index/announcements/qwen-3-8-max-and-deepseek-v4-flash-added/): Alibaba's Qwen 3.8 Max and DeepSeek v4 Flash 0731 join the Code Index, landing at opposite ends of the cost spectrum.- [GPT 5.6 Luna and Terra added!](https://research.revelo.com/code-index/announcements/new-models-added/): GPT 5.6 Terra and GPT 5.6 Luna join revision 1 of the Code Index, entering the overall top five and redefining the cost-efficiency frontier.- [Frontier-Bench is now Terminal Bench 3.0](https://research.revelo.com/code-index/announcements/renaming-frontier-bench-to-terminalbench/): Frontier-Bench has a new name and our benchmark is following suit.- [Code Index is live!](https://research.revelo.com/code-index/announcements/big-bang-release/): The first published revision of Revelo's Code Index is now available. ## About - [About](https://research.revelo.com/code-index/about/): Why Revelo builds code datasets, evaluations, and the Code Index. - [Revelo](https://www.revelo.com/): The company behind the research.