Qwen opens Qwen3.8-27B weights, touts 262K context

Qwen has just rolled out Qwen3.8-27B, an Apache 2.0-licensed 27B multimodal model with a native 262K-token context window. The company says it beats Qwen3.7-Plus, and early testers are already pushing local builds and quantized runs.

qwen cover

TL;DR

  • Qwen3.8-27B open weights: 27B-parameter native multimodal dense model, Apache 2.0 licensed
  • Context window: Native 262K tokens, extendable to 1M tokens via YaRN
  • Distribution: Available on Hugging Face and ModelScope
  • Positioning: Qwen3.8-2.4T-A95B for agent-building; 27B aimed at local applications
  • Qwen benchmark claims: 73.0 terminal coding, 79.0 software engineering, 90.3 competitive coding, 91.1 document intelligence
  • Local deployment reports: Unsloth Dynamic GGUFs and Unsloth Desktop support; quantized run on dual RTX 3090 at 39 tok/s via llama.cpp

Qwen has released the open weights for Qwen3.8-27B, a 27-billion-parameter native multimodal dense model licensed under Apache 2.0. The company claims it outperforms Qwen3.7-Plus overall, with a particular focus on coding and office workflows.

Qwen3.8-27B supports a native 262K-token context window, which Qwen reports can be extended to 1 million tokens through YaRN. The release is available through Hugging Face and ModelScope.

Qwen also points to open weights for the Qwen3.8-2.4T-A95B model, described in the announcement as “Max-level.” That model is positioned for agent-building, while the smaller 27B release is intended for local applications.

Qwen’s benchmark claims

A benchmark table published alongside the release compares Qwen3.8-27B with Qwen3.6-27B, Qwen3.7-Plus, Muse Glimmer-30B and Opus4.6 Max. Qwen3.8-27B records scores of 73.0 on agentic terminal coding, 61.7 on agentic coding, 42.3 on repo-level code generation and 79.0 on software engineering.

The same table lists scores of 70.7 on long-horizon office work, 79.5 on instruction following, 89.2 on scientific reasoning and 90.3 on competitive coding. On multimodal tasks, the model scores 84.3 on computer use, 64.8 on browser use, 81.9 on mobile use and 91.1 on document intelligence.

Those figures come from Qwen’s own evaluation material, and the supplied table marks the best result in each row in bold. They should therefore be treated as release claims rather than independent verification of the model’s overall performance.

Local deployment draws immediate interest

Unsloth claims its Dynamic GGUFs allow Qwen3.8-27B to run locally, adding that the model is supported for running and fine-tuning in Unsloth Desktop. Several replies to Qwen’s announcement focused on hardware requirements, including whether the model could run on systems with 8GB of VRAM, 24GB of memory or 64GB of RAM.

One user reported running a quantized Qwen3.8-27B build on a dual-RTX 3090 Linux system at 39 tokens per second, using llama.cpp and a 262,144-token context setting. That is a single user report rather than a general hardware recommendation.

Another post described creating a shooting game with Qwen3.8, Ollama and Grok Build on an Alienware system with an RTX 4090. The reports indicate early experimentation with local deployment, while questions about minimum hardware and laptop support remain prominent in the discussion.

Source: Qwen on X

Continue the conversation on Slack

Did this article spark your interest? Join our community of experts and enthusiasts to dive deeper, ask questions, and share your ideas.

Join our community