Announcement

GPT 5.6 Luna and Terra added!

GPT 5.6 Terra and GPT 5.6 Luna join revision 1 of the Code Index, entering the overall top five and redefining the cost-efficiency frontier.

Our Code Index has new models, and they make an immediate mark.

OpenAI’s GPT 5.6 Terra and GPT 5.6 Luna have both landed in the sixth and seventh position respectively.

GPT 5.6 Luna is simply the cheapest model we have evaluated, at $0.09 per task on average. That puts it below low-cost open-weight models such as Kimi K3 ($1.12) and Tencent’s Hy3 ($0.25), and it also arrives at a solution much faster: 6.6 minutes per task on average, against 39.6 and 43.8 minutes respectively.

GPT 5.6 Terra does not sit too far from it. It picks up a 0.9-point increment in the overall Index (37.1 against 36.2), while keeping the average cost per task under one dollar, at $0.54.

These two now sit on the cost-efficiency frontier alongside GPT 5.6 Sol, Claude Opus 5, and Kimi K3, meaning no evaluated model beats them on both price and score.

Although both cheap and efficient, when put side-by-side, GPT 5.6 Luna presents itself as the most interesting option. Almost the same success-rate, but for one-sixth the price.

Efficiency charts plotting pass rate against mean agent time and mean cost per task, with GPT 5.6 Terra and GPT 5.6 Luna highlighted on the cost-efficiency frontier
Pass rate against mean agent time and mean cost per task, with the cost-efficiency frontier highlighted. Explore the efficiency charts.

Both models are already reflected on the live board. See the leaderboard.

← Back to Announcements

Talk to us

Revelo works with model labs on:

  • Training datasets for code generation
  • Evaluations of new and unreleased models
  • Leaderboard analysis, including failure modes and agent trajectories

Talk to us

Tell us whether you need a dataset, an evaluation, or help analyzing model results.

Loading form…