Benchmark
SWE-Lancer
These are freelance engineering jobs carried over with their original prices. Each task keeps the client's specification and the amount that was actually paid for the work. The price gives the pass rate some economic context, though the Code Index still weights every task equally.
The format comes from OpenAI's SWE-Lancer, published in February 2025 with 1,488 Upwork tasks worth $1 million between them. Ours are different jobs, posted on Freelancer, Upwork, RedditHire, and 99Lancers, and we kept them private until evaluation. Each one carries the closing price of the job it came from, and it passes only when every required test succeeds. There are 20 tasks in the set.
Abbreviated example task
Build "a simple but modern PyQt5 UI" for a client. The application needs login and bot windows, registration, form validation, password rules, light and dark themes, and a window orchestrator. The original specification includes a screenshot. The verifier runs 91 offscreen Qt tests, and every required test must pass.Task mix
The domain and empirical difficulty of the 20 tasks in this set.
Domain spread
tasks per domain
Difficulty
Challenging: 0% pass rate. Hard: above 0% to 33%. Medium: above 33% to 66%. Easy: above 66% to below 100%. Trivial: 100%.
tasks per difficulty
Talk to us
Revelo works with model labs on:
- Training datasets for code generation
- Evaluations of new and unreleased models
- Leaderboard analysis, including failure modes and agent trajectories