Alibaba’s artificial intelligence model Qwen3.8-Max successfully rose to first place in the Code Arena: WebDev ranking of the Arena.ai platform. The developers did not release a new version, but conducted additional intensive training on it, focused on long-term work on programming and agents, and launched the Qwen3.8-Max-0902 version.
Image source: Alibaba Group.com
Model ID Qwen3.8-Max-0902 indicates that the upgrade has not yet been performed. The developers made a small update to the model, essentially keeping the original metrics: 2.4 trillion parameters, 95 billion active parameters per token, and prices of $2 per 1 million input tokens and $6 per 1 million output tokens.
On August 3, the first version of Alibaba Qwen3.8-Max was released. The model ranked fourth in the Code Arena: WebDev score, scoring 1668 points. It lost to the Anthropic Claude Opus 5 (Max) (1705 points), Moonshot Kimi K3 (Max) (1676 points) and Claude Opus 5 (High) (1669 points). Repeated tests on September 2 changed the balance of power: according to the new results, Qwen3.8-Max-0902 scored 1691 points, 2 points ahead of Claude Opus 5 (Max) and 17 points ahead of Kimi K3 (Max). The model maintained its strong position across all competitions, taking first place in the categories “Data and Analytics” and “Consumer Products”; second place in the “Branding and Marketing”, “Games” and “Simulation” categories; and third place in the Content Creation Tools and Reference Design categories.
Alibaba achieved these results through targeted post-training in two areas: programming and “collaboration”—collaborative tasks between agents with a long-term horizon. In such tasks, the model coordinates the multi-stage work of artificial intelligence subagents during operations using documents, interfaces, and code files. In the TerminalBench 3.0 test, which measures the model’s ability to perform multi-step coding work in a real terminal environment, the score increased from 11.3 points to 29.0 points, an increase of 2.6 times. In ProgramBench, the improvement increased from 10.5 points to 28.0 points. In the JobBench test of office automation, the improvement increased from 53.4 points to 64.0 points. In the WorkArena Elo test for multi-step web page navigation and task completion, results jumped from 1348 to 1468.
Judging from Alibaba’s own testing, Qwen3.8-Max-0902 is inferior to Claude Opus 5 in TerminalBench, DeepSWE, NL2Repo, ProgramBench, SWE-Marathon, CoWorkBench and Toolathlon benchmark tests. The first version of the model, released on August 3, showed a score of 56.6 on the DeepSWE test of agent-based long-term coding skills – Gemini (70 points) and OpenAI GPT-5.6 Sol (73 points) were also higher; for Qwen3.8-Max-0902, updated data for this test has not yet been released. The first iteration of Qwen3.8-Max scored 67.7 points in the SWE-bench Pro test of performing real engineering tasks on the GitHub platform, with Claude Fable 5 scoring 80 points. The Hugging Face platform is now live with the open weight Alibaba Qwen 3.8-2.4T-A95B, released on August 12th; plans to release them for the index 0902 variant have not yet been announced.
If you find an error, select it with your mouse and press CTRL+ENTER.










