DeepSeek Releases V4-Flash-0731 API in Public Beta

DeepSeek placed the official V4-Flash API into public beta on July 31, updating the served model to V4-Flash-0731 while keeping the existing API call unchanged. Its documentation lists $0.14 per million uncached input tokens and $0.28 per million output tokens, below OpenAI's newly reduced GPT-5.6 Luna rates of $0.20 and $1.20.
DeepSeek placed the official V4-Flash API into public beta on July 31 and updated the served model to DeepSeek-V4-Flash-0731. The company says the API call remains unchanged: developers continue to use the deepseek-v4-flash model name.
The update follows the V4 preview released on April 24. DeepSeek's changelog says the 0731 version keeps the preview model's architecture and size but received new post-training. The change applies to the V4-Flash API; the V4-Pro API and DeepSeek's app and web models were not included.
What changed
DeepSeek says V4-Flash-0731 has stronger agent capabilities and now supports the Responses API natively, with specific adaptation for Codex. Its published results include 82.7 on Terminal Bench 2.1, 54.4 on DeepSWE and 70.3 on Toolathlon Verified. These are vendor-reported benchmarks, not independent comparisons, and two of the listed test sets are internal.
The company's current documentation lists a 1 million-token context window, maximum output of 384,000 tokens, tool calling, JSON output and both thinking and non-thinking modes. Developers should test those features against their own prompts and tool chains rather than infer production reliability from benchmark scores alone.
The pricing comparison
DeepSeek's pricing page lists V4-Flash at $0.14 per million uncached input tokens, $0.0028 per million cache-hit input tokens and $0.28 per million output tokens. It also says a future peak-hours policy will double prices during specified Beijing-time windows, but that policy's effective date remains subject to a later announcement.
OpenAI's July 30 update reduced GPT-5.6 Luna by 80%, to $0.20 per million input tokens and $1.20 per million output tokens. On headline uncached rates, DeepSeek remains cheaper on both sides of the request. A simple one-million-input plus one-million-output comparison is $0.42 for V4-Flash and $1.40 for Luna, but actual bills depend on prompt-to-output ratios, caching, retries and tool-call behavior.
What teams should measure
The price spread makes V4-Flash relevant for high-volume coding and agent workflows, but token price is only one part of deployment cost. Teams considering a switch should measure task success, output length, latency, concurrency, tool-call accuracy, failure rates, data controls and regional availability on the workloads they actually run.
DeepSeek's official release establishes the product change and current rates. It does not independently validate the company's benchmark claims or establish that V4-Flash is interchangeable with Luna for a particular production system.
Key Points
- 1DeepSeek's July 31 update moves the official V4-Flash API into public beta and serves the post-trained V4-Flash-0731 model through the existing model name.
- 2Official pricing lists $0.14 per million uncached input tokens and $0.28 per million output tokens, below GPT-5.6 Luna's newly reduced $0.20 and $1.20 rates.
- 3DeepSeek reports stronger agent benchmarks and native Responses API support, but production decisions still require workload-specific quality, latency, reliability and governance testing.
Scoring Rationale
The official V4-Flash-0731 API update combines a new post-trained model with very low published rates and native Responses API support. It is materially relevant to high-volume agent and coding workloads, though vendor benchmarks require independent workload testing.
Sources
Primary source and supporting public references used for this report.
Practice interview problems based on real data
1,625 SQL & Python problems across 15 industry datasets — the exact type of data you work with.
Try 250 free problems

