These tasks reach past everyday coding into software engineering, business operations, science, machine learning, and security. What the agent hands in changes with the work: a patch, a trained checkpoint, a data artifact, a written analysis. It works inside an isolated container and is graded on whatever it leaves behind.

The format is Frontier-Bench, built by the Terminal-Bench and Harbor team as a harder and broader successor to Terminal-Bench. It keeps that benchmark's split between the agent container and the verifier container. Artifacts are pulled out when the run ends and graded somewhere clean, which shuts the door on reward hacking and means a task can be re-graded later if its verifier changes. There are 20 tasks in the set, across 5 domains.

Abbreviated example task

Build an exact polynomial normal-form engine that reconstructs a finite-rank commutative tower quotient algebra over the rationals from legacy transcript rows — selecting the active rows, ordering the monic border relations into a triangular tower, and resolving affine alias-of-alias chains — before any expression can be reduced, and it is graded byte-for-byte against a sealed reducer on profiles you never see.

Task mix

The domain and empirical difficulty of the 20 tasks in this set.

Domain spread

software-engineering 8
business-operations 6
science 3
ml 2
security 1

tasks per domain

Difficulty

Challenging: 0% pass rate. Hard: above 0% to 33%. Medium: above 33% to 66%. Easy: above 66% to below 100%. Trivial: 100%.

challenging 8
hard 10
medium 2
easy 0
trivial 0

tasks per difficulty

See Frontier-Bench results →

Talk to us

Revelo works with model labs on:

  • Training datasets for code generation
  • Evaluations of new and unreleased models
  • Leaderboard analysis, including failure modes and agent trajectories

Talk to us

Tell us whether you need a dataset, an evaluation, or help analyzing model results.

Loading form…