DFlash 2 touts 70 token/sec Qwen3.8-27B on M5 Max

DFlash 2 claims it can run Qwen3.8-27B at up to 70 tokens per second on an M5 Max MacBook Pro, with “up to 4.6×” faster decoding at the same output. Early user reports vary widely, and cross-hardware comparisons remain unclear.

qwen cover

TL;DR

  • DFlash 2 performance claim: Qwen3.8-27B at 70 tokens/s on M5 Max MacBook Pro
  • Speedup claim: “Up to 4.6×” faster than autoregressive decoding while producing the same output
  • Update behavior: Adds one accepted token per pass; originated at Z Lab, upgraded at Inco AI
  • New drafters: For Qwen3.8-27B and Meta’s Muse Glimmer; compatible with SGLang, vLLM, llama.cpp, oMLX
  • Technique notes: Lightweight selector traces coherent path through parallel candidates; two-tap convolution targets end-of-block quality
  • Reported variability: 42.5 tokens/s on DGX Spark; 53 tokens/s; 28 tokens/s on oMLX with think mode disabled

DFlash 2 claims to run Qwen3.8-27B at 70 tokens per second on an M5 Max MacBook Pro, according to an announcement from Zhijian Liu. Liu claims the system can reach “up to 4.6×” the speed of autoregressive decoding while producing the same output.

The release is the latest version of DFlash, which Liu describes as having been seeded at Z Lab and upgraded at Inco AI. The update reportedly adds one accepted token on every pass. DFlash models have surpassed 3.5 million downloads on Hugging Face, with the Qwen3.6 drafters accounting for most of that total, according to Liu.

Two new drafters

DFlash 2 launches with two drafters: one designed for Qwen3.8-27B and another for Meta’s Muse Glimmer. Both are listed as compatible with SGLang, vLLM, llama.cpp, and oMLX.

Liu describes the system as using a lightweight selector to trace a coherent path through parallel candidates generated by the drafter. A two-tap convolution is intended to prevent the draft from losing quality toward the end of each block.

The release also serves as an early look at Inco AI, which Liu mentions is developing an end-to-end inference stack for agent workloads. Further details about the company were not included in the announcement.

Reported results vary by setup

The 70-token-per-second result is not repeated across every reported test. One user reported 42.5 tokens per second on a DGX Spark, up from 20 tokens per second, while another reported 53 tokens per second. A separate oMLX user reported 28 tokens per second with think mode disabled.

Those figures do not establish a direct comparison, since the posts do not provide a common hardware and configuration setup. Liu requested configuration details from the user reporting 53 tokens per second, but no adjustment or explanation was included in the supplied material.

A separate benchmark shared in a quoted post evaluated a Qwen3.8-27B NVFP4 quantization from RadixArk using SGLang and RadixArk’s DSpark speculator. On a single 600-watt NVIDIA RTX PRO 6000 Blackwell Workstation Edition GPU, the test recorded 4,174.23 tokens per second for prompt processing and 189.32 tokens per second for token generation, with a peak generation rate of 266 tokens per second.

That benchmark used a 32,768-token context, a concurrency of one, and a single generation run with llama-benchy v0.4.0 on August 16, 2026. Its hardware and test conditions differ substantially from the M5 Max result, so the figures should not be treated as like-for-like measurements.

Source: Zhijian Liu on X

Continue the conversation on Slack

Did this article spark your interest? Join our community of experts and enthusiasts to dive deeper, ask questions, and share your ideas.

Join our community