Kimi launches K3 AI model with 1M-token context

Kimi.ai has just rolled out Kimi K3, touting 2.8 trillion parameters, a 1 million-token context window, and native multimodal support. The launch spotlights new attention upgrades and MoE scaling claims, with open weights slated for July 27, 2026.

kimi cover

TL;DR

  • Kimi K3: Claimed 2.8T parameters, 1M-token context, native multimodal support
  • Availability: Live on Kimi.ai, Kimi Work, Kimi Code, and Kimi API; open weights planned July 27, 2026
  • Architecture updates: Kimi Delta Attention (KDA), Attention Residuals (AttnRes), Stable LatentMoE activating 16/896 experts
  • Efficiency claims: ~2.5× scaling efficiency vs K2; KDA up to 6.3× faster decoding at 1M tokens
  • Training/infra claims: AttnRes ~25% higher training efficiency; kernel iteration cut FP+BP 283.6→114.4 ms
  • Benchmarks: Coding Terminal Bench 2.1 88.3; leads Program Bench 77.8, SWE Marathon 42.0; mixed vs Fable 5

Kimi.ai has introduced Kimi K3, a model the company describes as “Open Frontier Intelligence,” with claims of 2.8 trillion parameters, a 1 million-token context window and native multimodal support. Kimi says the system is live on its website, Kimi Work, Kimi Code and the Kimi API, with open weights planned for July 27, 2026.

The launch materials place Kimi K3 on top of two architecture updates: Kimi Delta Attention (KDA) and Attention Residuals (AttnRes). Kimi claims those changes improve how information moves across sequence length and model depth, while a Stable LatentMoE setup activates “16 out of 896 experts.” The company also attributes an “approximate 2.5× improvement” in scaling efficiency versus K2 to the revised architecture, training recipe and data mix.

Kimi further asserts that Kimi Delta Attention can deliver “up to 6.3x faster decoding” in million-token contexts, while Attention Residuals provide “~25% higher training efficiency at <2% additional cost.” In a separate optimization example, the company says K3 spent 15 hours iterating on an AttnRes kernel at production scale and cut forward-plus-backward time from 283.6 ms to 114.4 ms without changing numerics.

The benchmark graphics released with the announcement paint a mixed picture. In coding, Kimi K3 reaches 88.3 on Terminal Bench 2.1, close to GPT-5.6 Sol’s 88.8, but trails Fable 5 on FrontierSWE and Kimi Code Bench 2.0. It leads on Program Bench with 77.8 and SWE Marathon with 42.0, though the margins are small in several cases. The charts are labeled with the note “All maxed out on thinking effort: max or xhigh.”

On general-agent tests, Kimi K3 posts 30.8 on Automation Bench and 91.2 on BrowseComp, while Fable 5 leads GDPval-AA v2 Elo, AA-Briefcase Elo and JobBench. In the visual-agent section, Kimi K3 scores 91.3 on CharXiv (RQ) with tool and 41.0 on Zerobench with tool, again behind Fable 5. Kimi’s internal knowledge-work benchmark claims are stronger: the company says Kimi K3 Max scored 75.5 on Online Exp Bench, 73.5 on DECK-Bench and 62.6 on Finance-Bench, ahead of Claude Opus 4.8 and GPT-5.5 in those internal evaluations.

The launch materials also emphasize “vision in the loop,” describing Kimi K3 as able to move between code and live screenshots while turning concepts, images and videos into interactive experiences. That pitch, along with the planned open-weight release, is likely to keep attention on how the model performs once independent testing becomes available.

Source: Kimi.ai on X

Continue the conversation on Slack

Did this article spark your interest? Join our community of experts and enthusiasts to dive deeper, ask questions, and share your ideas.

Join our community