Announcement
Qwen 3.8 Max and DeepSeek v4 Flash 0731 added!
Alibaba's Qwen 3.8 Max and DeepSeek v4 Flash 0731 join the Code Index, landing at opposite ends of the cost spectrum.
Alibaba’s Qwen 3.8 Max and DeepSeek v4 Flash 0731 were both released last weekend and are already part of our leaderboard!
The score of Qwen 3.8 Max (tested with Claude Code) is impressive, occupying 7th place overall and 2nd among the open-weight models. It also places itself in the frontier of time efficiency when compared only to the open-weight models, with an average of 39 min per task. However, this model comes with a cost, and it is huge. Qwen 3.8 Max is the most expensive one we evaluated so far, costing $8.93 per task ($1.16 more than Claude Fable 5, the second most expensive).
DeepSeek v4 Flash 0731 comes with an exact opposite scenario. This model is the 2nd cheapest among all models (the cheapest compared to the open-weight ones), at only $0.14 per task. This is why it is being vastly used, toping the most-used-models leaderboard of @OpenRouter right now. However, you will need to be patient, as this is the slowest model of all, taking ~45 minutes per task. Its score reached 30.4 (evalted with Terminus 2), which places it right after @Zai_org’s GLM 5.2.
Talk to us
Revelo works with model labs on:
- Training datasets for code generation
- Evaluations of new and unreleased models
- Leaderboard analysis, including failure modes and agent trajectories