DeepSeek puts nNew V4 Flash in public beta with agent upgrades

DeepSeek has just rolled out V4 Flash in public beta, with the official API now live and upgraded agent capabilities. Company-posted benchmarks show gains over prior DeepSeek preview builds, though several results still trail Opus-4.8.

deepseek cover

TL;DR

  • V4 Flash public beta: Official API now live; agent capabilities upgraded
  • API compatibility: Native Responses API support; fully adapted for Codex (see official API docs)
  • Benchmark gains vs prior DeepSeek previews: Higher across all listed measures than V4-Flash-Preview and V4-Pro-Preview
  • Reported scores: Terminal Bench 2.1 82.7, NL2Repo 54.2, DeepSWE 54.4, DSBench-Hard 59.6
  • Comparison vs Opus-4.8: Trails on several benchmarks, including Terminal Bench 2.1, NL2Repo, DeepSWE, DSBench-Hard
  • Methodology notes: Upcoming DeepSeek Harness “minimal mode”; max tier; topp=0.95; temperature=1.0; some benchmarks internal/company-reported

DeepSeek has put V4 Flash into public beta, according to a post the company quoted as announcing that the "official API is now LIVE" and that agent capabilities were "massively upgraded." DeepSeek also claims the model now natively supports the Responses API format and is "fully adapted for Codex," with configuration details linked in its official API docs.

The accompanying comparison places DeepSeek-V4-Flash-0731 against V4-Flash-Preview, V4-Pro-Preview, GLM-5.2 and Opus-4.8 across a set of code and agent benchmarks. In that table, the new version posts higher scores than DeepSeek-V4-Flash-Preview and DeepSeek-V4-Pro-Preview on every listed measure, including Terminal Bench 2.1 at 82.7, NL2Repo at 54.2, DeepSWE at 54.4 and DSBench-Hard at 59.6. It also trails Opus-4.8 in several rows, including Terminal Bench 2.1, NL2Repo, Cybergym, DeepSWE, Toolathlon-Verified, Agents’ Last Exam, AutomationBench (Public), DSBench-FullStack and DSBench-Hard.

Fine print below the table states that public Code Agent tasks were tested with an upcoming DeepSeek Harness "minimal mode" framework using max tier, topp=0.95 and temperature=1.0. DeepSeek also describes DSBench-FullStack and DSBench-Hard as internal benchmark sets, so the numbers should be treated as company-reported results rather than independently verified comparisons.

Source: DeepSeek AI

Continue the conversation on Slack

Did this article spark your interest? Join our community of experts and enthusiasts to dive deeper, ask questions, and share your ideas.

Join our community