Benchmark
Terminal Bench 2.1
These tasks are multi-step work in a live command line. They cover building, debugging, configuration, data processing, and system repair, and the agent does all of it through a terminal inside an isolated container.
The format is Terminal-Bench 2.1, maintained by the Terminal-Bench project. Tasks are written by hand into containers, the reference solutions are written by people, and the verifiers judge the final state of the machine rather than the commands that got there. We matched the domain proportions Terminal-Bench 2.0 uses. There are 20 tasks in the set.
Abbreviated example task
Repair an unfinished backup recovery utility in /app. Seven Python files must merge split archive directories, verify SHA-256 manifests, remove duplicates, and apply include, exclude, and rename rules. Inline issue notes identify the known bugs. The verifier replays merges and inspects the final files, so any implementation that produces the correct state can pass.Task mix
The domain and empirical difficulty of the 20 tasks in this set.
Domain spread
tasks per domain
Difficulty
Challenging: 0% pass rate. Hard: above 0% to 33%. Medium: above 33% to 66%. Easy: above 66% to below 100%. Trivial: 100%.
tasks per difficulty
See Terminal Bench 2.1 results →
Talk to us
Revelo works with model labs on:
- Training datasets for code generation
- Evaluations of new and unreleased models
- Leaderboard analysis, including failure modes and agent trajectories