DFlash 2 claims to run Qwen3.8-27B at 70 tokens per second on an M5 Max MacBook Pro, according to an announcement from Zhijian Liu. Liu claims the system can reach “up to 4.6×” the speed of autoregressive decoding while producing the same output.
The release is the latest version of DFlash, which Liu describes as having been seeded at Z Lab and upgraded at Inco AI. The update reportedly adds one accepted token on every pass. DFlash models have surpassed 3.5 million downloads on Hugging Face, with the Qwen3.6 drafters accounting for most of that total, according to Liu.
Two new drafters
DFlash 2 launches with two drafters: one designed for Qwen3.8-27B and another for Meta’s Muse Glimmer. Both are listed as compatible with SGLang, vLLM, llama.cpp, and oMLX.
Liu describes the system as using a lightweight selector to trace a coherent path through parallel candidates generated by the drafter. A two-tap convolution is intended to prevent the draft from losing quality toward the end of each block.
The release also serves as an early look at Inco AI, which Liu mentions is developing an end-to-end inference stack for agent workloads. Further details about the company were not included in the announcement.
Reported results vary by setup
The 70-token-per-second result is not repeated across every reported test. One user reported 42.5 tokens per second on a DGX Spark, up from 20 tokens per second, while another reported 53 tokens per second. A separate oMLX user reported 28 tokens per second with think mode disabled.
Those figures do not establish a direct comparison, since the posts do not provide a common hardware and configuration setup. Liu requested configuration details from the user reporting 53 tokens per second, but no adjustment or explanation was included in the supplied material.
A separate benchmark shared in a quoted post evaluated a Qwen3.8-27B NVFP4 quantization from RadixArk using SGLang and RadixArk’s DSpark speculator. On a single 600-watt NVIDIA RTX PRO 6000 Blackwell Workstation Edition GPU, the test recorded 4,174.23 tokens per second for prompt processing and 189.32 tokens per second for token generation, with a peak generation rate of 266 tokens per second.
That benchmark used a 32,768-token context, a concurrency of one, and a single generation run with llama-benchy v0.4.0 on August 16, 2026. Its hardware and test conditions differ substantially from the M5 Max result, so the figures should not be treated as like-for-like measurements.
Source: Zhijian Liu on X



