Summary

On September 2, 2026, Alibaba released Qwen3.8-Max-0902, a post-training refresh of Qwen3.8-Max focused on coding and long-horizon agent work. Architecture (2.4T MoE, 1M-token context) and pricing ($2/$6 per million tokens) are unchanged, but every coding benchmark improved: Code Arena WebDev jumps to 1,691 (first overall, 3 points above Claude Opus 5 Max), TerminalBench 3.0 more than doubles from 11.3 to 29.0, and ProgramBench Almost Solved goes from 10.5 to 28.0.

What changed

Alibaba shipped Qwen3.8-Max-0902 as a snapshot update to the base Qwen3.8-Max model. Same 2.4T MoE architecture and 1M context, same $2/$6 API pricing; the delta is a post-training run that lifted all eight coding benchmarks, with the biggest gains on agentic terminal work (TerminalBench 3.0 11.3 → 29.0) and long-horizon program synthesis (ProgramBench Almost Solved 10.5 → 28.0). Qwen also flagged stronger native vision.

Why it matters

A same-price, drop-in checkpoint that more than doubles TerminalBench turns Qwen3.8-Max into a live-option coding agent model rather than a benchmark curiosity, especially for teams that already route to it through Vercel AI Gateway or self-host from the open weights. Alibaba proving it can move a frontier-scale model up on coding purely via post-training — while Anthropic and OpenAI charge more for their new coding SKUs — puts fresh pressure on Western pricing for agentic engineering work.

Evidence excerpt

Qwen3.8-Max-0902 ranks first overall on Code Arena WebDev with 1,691 points, 3 points above Claude Opus 5 Max (1,687). The largest gains are on TerminalBench 3.0 (11.3 to 29.0) and ProgramBench Almost Solved (10.5 to 28.0), both more than doubling. All 8 coding benchmarks improved. The architecture is unchanged: 2.4 trillion parameters, 1M token context, and the same pricing ($2/1M input, $6/1M output).

Sources