These tasks are multi-step work in a live command line. They cover building, debugging, configuration, data processing, and system repair, and the agent does all of it through a terminal inside an isolated container.

The format is Terminal-Bench 2.1, maintained by the Terminal-Bench project. Tasks are written by hand into containers, the reference solutions are written by people, and the verifiers judge the final state of the machine rather than the commands that got there. We matched the domain proportions Terminal-Bench 2.0 uses. There are 20 tasks in the set.

Abbreviated example task

Repair an unfinished backup recovery utility in /app. Seven Python files must merge split archive directories, verify SHA-256 manifests, remove duplicates, and apply include, exclude, and rename rules. Inline issue notes identify the known bugs. The verifier replays merges and inspects the final files, so any implementation that produces the correct state can pass.

Task mix

The domain and empirical difficulty of the 20 tasks in this set.

Domain spread

software-engineering 5
data-science 4
personal-assistant 2
security 2
data-processing 1
debugging 1
file-operations 1
machine-learning 1
model-training 1
scientific-computing 1
system-administration 1

tasks per domain

Difficulty

Challenging: 0% pass rate. Hard: above 0% to 33%. Medium: above 33% to 66%. Easy: above 66% to below 100%. Trivial: 100%.

challenging 0
hard 8
medium 12
easy 0
trivial 0

tasks per difficulty

See Terminal Bench 2.1 results →

Talk to us

Revelo works with model labs on:

  • Training datasets for code generation
  • Evaluations of new and unreleased models
  • Leaderboard analysis, including failure modes and agent trajectories

Talk to us

Tell us whether you need a dataset, an evaluation, or help analyzing model results.

Loading form…