Alibaba's Qwen3.8-Max packs 2.4 trillion parameters and undercuts Western rivals on price

Alibaba has released Qwen3.8-Max, a 2.4 trillion-parameter mixture-of-experts model that the company says is the most capable in its Qwen family to date. The launch pushed Alibaba's Hong Kong shares up as much as 7% and its US-listed shares roughly 4.5%, as investors reacted to a model that undercuts Western frontier labs on price while posting competitive benchmark scores.
What the numbers actually show
Qwen3.8-Max is a mixture-of-experts architecture, meaning only a fraction of its 2.4 trillion total parameters activate for any given query — Alibaba has not disclosed the exact activated-parameter count, following a pattern common among Chinese labs of withholding that detail. The model accepts text, image, and video input and returns text, with a context window supporting roughly 991,000 tokens of input and 131,000 tokens of output — enough to process an entire codebase or a lengthy legal document in a single pass.
On Terminal-Bench 2.1, a benchmark measuring real-world coding and agentic task performance, Qwen3.8-Max scored 86.6, ahead of Anthropic's Claude Opus 4.8 at 84.6, though it trails OpenAI's GPT-5.6 Sol at 88.8. The gains are more dramatic on DeepSWE 1.1, an agentic software engineering benchmark, where the model jumped from 21.6 to 56.6 compared to its predecessor — more than doubling its score. Multimodal performance is also strong: 86.1 on OSWorld-Verified and 92.1 on OmniDocBench 1.5, a document understanding benchmark.
The pricing angle
Alibaba is pricing the hosted API through DashScope at $2.00 per million input tokens and $6.00 per million output tokens, with cached input tokens available at just $0.25 per million — roughly eight times cheaper than fresh input. That pricing sits well below what US frontier labs typically charge for comparably capable models, continuing a pattern where Chinese AI labs compete aggressively on cost even when they trail slightly on raw benchmark performance.
Deployment reality check
The full 2.4 trillion-parameter checkpoint requires multi-node datacenter infrastructure to run, putting genuine self-hosted deployment out of reach for all but the largest AI infrastructure operators. For teams wanting to run the model on standard GPU hardware, Alibaba is also offering Qwen3.8-27B, a smaller dense model from the same family, as the realistic on-premise option. Open weights for both the 2.4T flagship and the 27B model were promised for release the week following this announcement.
Where this fits in the broader race
On the crowdsourced Arena.AI leaderboard, Qwen3.8-Max immediately became the top-ranking Chinese model for text tasks, though it still trails several Anthropic offerings including Claude Fable 5. For vision tasks specifically, it ranked second globally, behind only a variant of Fable 5 — a notable result for a model whose training and inference costs are reported to run a fraction of what comparable Western models cost to serve.
As reported by MarkTechPost, the release lands just weeks after rival Chinese lab Moonshot AI shipped its open-weight Kimi K3 model, underscoring how quickly the competitive cycle between Chinese AI labs has compressed in 2026.
Originally reported by MarkTechPost. Read the original article for additional details.
View original source