Announcement

GPT 6 Astra, DeepSeek V4.1 Flash and Gemini 3.8 Flash added

Three new models join the Code Index, and SWE-Bench-Pro Hard leaves the index until we re-run it without internet access.

Three new models are on the Code Index: OpenAI’s GPT 6 Astra, DeepSeek V4.1 Flash and Google’s Gemini 3.8 Flash. With this update, SWE-Bench-Pro Hard no longer counts toward the index, so the index is now the mean of SWE-Lancer, Terminal Bench 2.1 and Terminal Bench 3.0.

The new models

  • GPT 6 Astra (Codex) is second, at 45.0. It is 3.3 points behind Claude Opus 5 and has the second-best Terminal Bench 3.0 score (30.0, against 31.7 for Claude Opus 5). It costs $4.99 per task on average, less than Claude Opus 5 ($7.15), Claude Fable 5 ($8.42) and Qwen3.8 Max ($8.55).
  • DeepSeek V4.1 Flash (DeepSeek Harness Minimal) is fifth, at 38.9, 0.5 points behind GPT 5.6 Sol. It ties GPT 6 Astra on Terminal Bench 2.1 (65.0). At $0.48 per task at DeepSeek list prices, it is the cheapest of the eight highest-ranked models, but it takes 30 minutes per task on average.
  • Gemini 3.8 Flash (Terminus 2) is sixth, at 36.7, 12.3 points above Gemini 3.6 Flash with the same agent. It shares the best SWE-Lancer score with Claude Fable 5 (51.7), but it scores 8.3 on Terminal Bench 3.0.

All three ran with each task’s own agent time limit (1×). Most models already on the board had more time; see the time limits below.

See the leaderboard.

SWE-Bench-Pro Hard is out of the index for now

Since September 2026, SWE-Bench-Pro Hard does not count toward the Code Index. The agents had internet access during these runs, and the fixes for these tasks are public in their upstream repositories, so a score can reflect finding the fix instead of solving the task. It will return to the index after we re-run it without internet access for all models. Its results stay on the site for reference.

What changed for the rest of the board

Every model that was already on the board has a lower index now, because SWE-Bench-Pro Hard was the easiest benchmark in the index. On the previous board, the models averaged 49.7% on SWE-Bench-Pro Hard, against 36.0% on Terminal Bench 2.1, 33.6% on SWE-Lancer and 10.7% on Terminal Bench 3.0. Claude Opus 5 stays first, at 48.3 (it was 57.1). The main rank changes among those models:

  • Claude Fable 5 (42.8) moves ahead of GPT 5.6 Sol (39.4), which had the best SWE-Bench-Pro Hard score on the previous board (86.7).
  • Kimi K3 (36.1) moves ahead of Claude Opus 4.8 (35.6).
  • GPT 5.6 Terra and GPT 5.6 Luna (both 30.6) move ahead of Qwen3.8 Max (29.4).
  • DeepSeek V4 Pro (17.8) and Hy3 are no longer tied; Hy3 and MiniMax M3 now share 15.0.

Agent time limits

The Methodology page now has a table of agent time limits. Each task sets its own agent time limit, and each run multiplies it by a factor. The factor was not the same for every run: the three new models, GPT 5.6 Terra, GPT 5.6 Luna, Qwen3.8 Max and DeepSeek V4 Flash 0731 ran at 1× on every benchmark, while the other models ran most of their trials at 2× or 3×.

Re-runs

For the three new models, we re-ran a trial only when it ended on an error from our infrastructure or the model provider, whatever its verifier result. Failures caused by the model were kept, and an agent timeout counted as a failure of the model, unless the agent never reached the model. This covered 32 DeepSeek V4.1 Flash trials that ended on provider errors, 10 of them provider content-filter blocks; 9 GPT 6 Astra trials on Terminal Bench 2.1 that timed out because the agent could not reach the model (three task images had no CA certificate bundle); and 1 Gemini 3.8 Flash trial that hit a sandbox command-length limit.

← Back to Announcements

Talk to us

Revelo works with model labs on:

  • Training datasets for code generation
  • Evaluations of new and unreleased models
  • Leaderboard analysis, including failure modes and agent trajectories

Talk to us

Tell us whether you need a dataset, an evaluation, or help analyzing model results.

Loading form…