DeepSeek's recent moves keep pointing at the same lever: cost per unit of useful work. On July 31 it released DeepSeek-V4-Flash-0731 as the official version of V4 Flash, superseding the preview while keeping its architecture and size, adding native Responses API support, and touching only the Flash API, leaving V4 Pro and the app and web models unchanged. Artificial Analysis scored Flash-0731 at 50 on its Intelligence Index, up from 40 for the previous V4 Flash, and reported its agentic-work Elo on GDPval-AA v2 rising from 1,189 to 1,559 with an unchanged 284-billion-total, 13-billion-active configuration. DeepSeek's own model card reports 82.7 on Terminal Bench 2.1, 54.2 on NL2Repo, 76.7 on Cybergym, 54.4 on DeepSWE and 70.3 on Toolathlon-Verified, but says those public code-agent benchmarks used an unreleased minimal mode of DeepSeek Harness at max reasoning effort, and identifies DSBench-FullStack and DSBench-Hard as internal test sets, which makes the figures vendor evidence rather than a substitute for your own workload testing. That lands just after the pricing change announced June 30, when DeepSeek said the official V4 would ship in mid-July and, for the first time, meter the API with peak and off-peak rates, doubling the cost of calls made between 9:00 a.m. and 12:00 p.m. and between 2:00 p.m. and 6:00 p.m. local time while off-peak prices stay flat. Together those two changes reward a specific engineering habit: move deferrable batch and evaluation work into off-peak windows, and re-benchmark with version pinning before assuming the Flash tier still behaves the way it did in preview. The demand side is moving in the same direction. UBS analysts led by Karl Keirstead found in June that roughly 60% of enterprises are throttling AI spending in some way and are using model routing to push routine tasks toward cheaper or open models, and CNBC reported that the share of tokens routed to Chinese models on OpenRouter has exceeded 30% every week since Feb. 8 and peaked at 46%, against a 12-month average of 11%.
The second thread is harder to price into a routing decision: DeepSeek is now a fixture in threat-intelligence reporting, and its supply of compute and capital is still unsettled. Palo Alto Networks' Unit 42 reported on July 30 that a Chinese-speaking operator drove DeepSeek through Hermes Agent to automate reconnaissance and unsuccessful Langflow and n8n exploitation attempts inside a campaign against more than 460 targets, with the three confirmed compromises traced to separate manual NetScaler exploitation, and its report contains an unresolved inconsistency involving claimed Marimo command execution. Hunt.io said on July 14 that an exposed operator directory showed Claude Code and DeepSeek-v4-pro wired into a suspected China-linked intrusion workflow, attributing confirmed compromises in Afghanistan, Thailand and Taiwan while describing U.S. activity as reconnaissance and phishing preparation; that is one firm's technical investigation, not independent attribution. Check Point reported on a real DeepSeek-generated sample, InfernoGrabber v9.0, in which V4 refused prompts using the word ransomware but produced the same functional browser-native encryption code under neutral phrasing. None of that describes a defect a deploying team can patch, but it does argue for correlating agent runtime, model endpoint, shell execution and credential telemetry rather than treating model traffic as a detection signal on its own. Meanwhile the constraints DeepSeek is trying to escape are visible in three July reports: Reuters said on July 7 that it has spent about a year quietly building its own inference chip to reduce dependence on Nvidia and Huawei; Beijing is reportedly weighing limited Nvidia H200 access for Alibaba, ByteDance and DeepSeek, possibly fewer than 200,000 chips and less than half of what was requested; and after closing a first external round of more than $7 billion at a valuation above $50 billion, the company was reported on July 14 to be in preliminary talks on a second financing at roughly $71 billion pre-money, which it has not confirmed.