Alibaba launches Qwen3.8-Max, a 2.4T-parameter AI model

Alibaba has just rolled out Qwen3.8-Max, a 2.4 trillion-parameter model aimed at coding, workplace tasks, and multimodal agent workflows. Qwen claims it can run autonomously for days on long projects, with open weights promised next week alongside Qwen3.8-27B.

qwen cover

TL;DR

  • Qwen3.8-Max (2.4T parameters): Positioned for coding, workplace tasks, and multimodal agentic workloads
  • Long-horizon autonomy claims: >10-day runs; cited 16-day coding trace, 500+ chip-design turns, 365-day e-commerce strategy
  • Multimodal loop: Vision used for planning, execution, and self-correction across extended tasks
  • Benchmarks (Qwen-published): TerminalBench-2.1 86.6; SWE-Pro 67.7; PaperBench 93.0; FrontierSWE 73.5
  • Additional scores: OSWorld-Verified 86.1; WebArena-Verified 66.8; AndroidWorld 85.3; MMMU-Pro 82.3; VideoMMLU 90.4
  • API pricing & access: Input $2/M, output $6/M, caching $0.25/M; available via Qwen Studio and API; open weights planned next week with Qwen3.8-27B

Qwen has introduced Qwen3.8-Max, a 2.4 trillion-parameter model that Alibaba’s Qwen team describes as its most capable release to date. The company positions it for coding, general workplace tasks and multimodal agentic workloads, while promising open weights for Qwen3.8-Max the following week alongside an open-weight Qwen3.8-27B model.

A model aimed at long-running tasks

Qwen claims Qwen3.8-Max can operate autonomously for more than 10 days, taking a software project from an empty folder to production without continuous human guidance. Another promotional post describes a 16-day autonomous coding run and links to a project trace on GitHub.

The company also claims the model completed more than 500 turns of chip-design optimization and managed a 365-day e-commerce strategy. Qwen presents these examples as evidence of long-horizon planning, adaptive learning and closed-loop execution rather than short, isolated question-and-answer tasks.

For multimodal work, Qwen describes vision as a continuous feedback loop used for planning, execution and self-correction. The company also claims production-quality deliverables across hundreds of professions, although the announcement does not define how those deliverables were evaluated.

Published benchmark results

Qwen’s published comparisons list Qwen3.8-Max at 86.6 on TerminalBench-2.1, 67.7 on SWE-Pro, 93.0 on PaperBench and 73.5 on FrontierSWE. The model also recorded 80.7 on QwenSWEBench and 58.4 on QwenCoderBench.

Other listed results include:

  • 1,724 on QwenReactBench
  • 74.8 on CoWorkBench
  • 53.4 on JobBench
  • 52.4 on Agents’ Last Exam
  • 92.6 on GPQA Diamond
  • 82.8 on IFBench
  • 66.3 on LongBench v2
  • 86.1 on OSWorld-Verified
  • 66.8 on WebArena-Verified
  • 85.3 on AndroidWorld

The multimodal table lists scores of 82.3 on MMMU-Pro, 91.9 on LogicVista and 90.4 on VideoMMLU. Qwen also reports 91.5 on Parametric CAD Bench and 92.1 on OmniDocBench 1.5.

These figures come from Qwen’s own comparison material, and the announcement does not provide independent verification or full testing conditions. A Text Arena ranking dated Aug. 1 places Alibaba second among labs, with Qwen3.8-Max scoring 1,496 ± 10 and holding a listed model rank of five. The entry is marked “Preliminary,” while Anthropic ranks first among labs at 1,509 ± 6.

Pricing and open-weight plans

Qwen lists API pricing at:

  • Input: $2 per million tokens
  • Output: $6 per million tokens
  • Implicit caching: $0.25 per million tokens

The model is available through Qwen Studio and an API, with additional technical details in Qwen’s product blog.

The planned open-weight releases are likely to attract particular interest around Qwen3.8-27B. Unsloth AI welcomed the announcement and mentioned plans to create quantized versions. Several other replies requested a 35B variant, but Qwen has not announced one in the supplied material.

Source: Qwen on X

Continue the conversation on Slack

Did this article spark your interest? Join our community of experts and enthusiasts to dive deeper, ask questions, and share your ideas.

Join our community