These are freelance engineering jobs carried over with their original prices. Each task keeps the client's specification and the amount that was actually paid for the work. The price gives the pass rate some economic context, though the Code Index still weights every task equally.

The format comes from OpenAI's SWE-Lancer, published in February 2025 with 1,488 Upwork tasks worth $1 million between them. Ours are different jobs, posted on Freelancer, Upwork, RedditHire, and 99Lancers, and we kept them private until evaluation. Each one carries the closing price of the job it came from, and it passes only when every required test succeeds. There are 20 tasks in the set.

Abbreviated example task

Build "a simple but modern PyQt5 UI" for a client. The application needs login and bot windows, registration, form validation, password rules, light and dark themes, and a window orchestrator. The original specification includes a screenshot. The verifier runs 91 offscreen Qt tests, and every required test must pass.

Task mix

The domain and empirical difficulty of the 20 tasks in this set.

Domain spread

software-engineering 20

tasks per domain

Difficulty

Challenging: 0% pass rate. Hard: above 0% to 33%. Medium: above 33% to 66%. Easy: above 66% to below 100%. Trivial: 100%.

challenging 0
hard 10
medium 9
easy 1
trivial 0

tasks per difficulty

See SWE-Lancer results →

Talk to us

Revelo works with model labs on:

  • Training datasets for code generation
  • Evaluations of new and unreleased models
  • Leaderboard analysis, including failure modes and agent trajectories

Talk to us

Tell us whether you need a dataset, an evaluation, or help analyzing model results.

Loading form…