DeepSeek-V4-Pro debuts with agent upgrades and new API pricing

DeepSeek has just rolled out DeepSeek-V4-Pro, bringing adjustable reasoning levels, stronger agent workflows, and native OpenAI Responses API support. It also introduces peak/off-peak API pricing starting Aug. 16, 2026, sparking debate over costs versus gains.

deepseek cover

TL;DR

  • Launched DeepSeek-V4-Pro for agent tasks, coding workflows, and tool use; model name deepseek-v4-pro
  • Availability: App/web “Expert Mode” and API; adds adjustable reasoning levels for V4-Pro and V4-Flash
  • OpenAI Responses API support and one-click setup optimized for Codex
  • Benchmarks (V4-Pro-0813): 87.9 Terminal Bench 2.1, 61.5 NL2Repo, 67.2 DSBench-Hard, 42.7/60 HLE
  • Evaluation setup: upcoming DeepSeek Harness minimal mode, max-tier; topp=0.95, temperature 1.0; results framework-dependent
  • API pricing changes: peak/off-peak rates; off-peak 50% lower; effective 16:00 UTC Aug 16, 2026; peak 01:00–04:00 and 06:00–10:00 UTC

DeepSeek has launched DeepSeek-V4-Pro, positioning the new model as an upgrade for agent tasks, coding workflows and tool use. The model is available through DeepSeek’s app and web interface in “Expert Mode,” as well as through its API, where the model name remains deepseek-v4-pro.

The release adds adjustable reasoning levels for both V4-Pro and V4-Flash. DeepSeek describes the options as low for simple tasks, high for regular agent workflows and max for complex tasks. The company also highlights native OpenAI Responses API support and a one-click setup optimized for Codex.

DeepSeek’s announcement focuses on “major Agent upgrades” and “strong production gains,” although it does not define the production baseline in the launch post. The company’s benchmark table compares V4-Pro-0813 with V4-Flash-0731, earlier V4 previews, GLM-5.2, Kimi-K3, Opus-4.8 and Fable 5 with fallback.

On the reported tests, V4-Pro-0813 scores 87.9 on Terminal Bench 2.1, compared with 82.7 for V4-Flash-0731 and 85.0 for Opus-4.8. It reaches 61.5 on NL2Repo, below Opus-4.8’s 69.7, and records 67.2 on DSBench-Hard, compared with 71.7 for Opus-4.8 and 68.3 for Fable 5. On HLE without tools, V4-Pro-0813 scores 42.7/60.0, while Opus-4.8 scores 49.8/57.9 and Fable 5 scores 53.3/63.0.

Other reported results include 83.3 on Cybergym, 62.7 on DeepSWE, 74.1 on Toolathlon-Verified and 31.8 on AutomationBench Public. V4-Flash-0731 scores 76.7, 54.4, 70.3 and 25.1 on those tests, respectively.

DeepSeek states that V4-Pro-0813 was evaluated on public Code Agent tasks using its upcoming DeepSeek Harness framework in minimal mode, with max-tier settings, topp=0.95 and a temperature of 1.0. It cautions that results may vary slightly when other frameworks are used, leaving the figures dependent on the testing setup.

API pricing changes

The release also brings a new pricing structure divided into peak and off-peak hours. DeepSeek claims that off-peak rates are 50% lower than peak rates. The new prices take effect at 16:00 UTC on August 16, 2026.

ModelPeriodInput cache hitInput cache missOutput
V4-FlashOff-peak$0.007/M$0.22/M$0.66/M
V4-FlashPeak$0.014/M$0.44/M$1.32/M
V4-ProOff-peak$0.022/M$0.66/M$1.98/M
V4-ProPeak$0.044/M$1.32/M$3.96/M

Peak hours run from 01:00–04:00 UTC and 06:00–10:00 UTC. All other hours are classified as off-peak.

The announced increase has prompted criticism from some users who viewed it as roughly a doubling of prices, particularly for V4-Pro. A separate comparison shared in response to the announcement argues that DeepSeek’s Pro pricing would remain below rates from several US-hosted providers even after a substantial increase. That comparison also suggests that V4-Flash is already priced closer to those services, though the figures come from an external social-media analysis rather than an independent pricing audit.

The difference between the two models is also drawing scrutiny. One reply questions whether V4-Pro offers enough of an advantage over Flash to justify its higher cost, pointing to relatively close benchmark results. The published figures do show a gap, but its size varies substantially by test: Pro leads Flash by 5.2 points on Terminal Bench 2.1, 7.3 points on NL2Repo and 7.6 points on DSBench-Hard.

DeepSeek has not announced open weights for V4-Pro. Unsloth AI, which previously supported V4-Flash, expressed interest in producing V4-Pro quantizations for local deployments.

Source: DeepSeek on X

Continue the conversation on Slack

Did this article spark your interest? Join our community of experts and enthusiasts to dive deeper, ask questions, and share your ideas.

Join our community