- The flagship model targets complex, long-context AI tasks with 384K-token output
- Pro charges up to 3X more than Flash, signaling a shift toward tiered pricing
DeepSeek officially launched V4 Pro, moving its flagship V4 series from preview to commercial deployment.
The model supports a 1-million-token context window—enough to process text roughly equivalent to the sci-fi novel The Three Body Problem trilogy—and a maximum output of 384,000 tokens, twice the previous limit.
V4 Pro ranks fourth globally on the Arena agent leaderboard and fifth on Aider’s coding leaderboard, matching its preview version on both benchmarks.
A tiered pricing strategy
V4 Pro is positioned as the premium, higher-end option, with pricing significantly above V4 Flash.
Per 1 million tokens, V4 Pro charges 0.025 yuan ($0.0037) for cache-hit input, 3 yuan for cache-miss input and 6 yuan for output, compared with 0.02 yuan, 1 yuan and 2 yuan, respectively, for V4 Flash.

Pro output tokens are therefore three times more expensive, while cache-miss input is also priced at three times the Flash rate.
The pricing split reflects a deliberate strategy: Flash targets high-volume, low-cost workloads, while Pro targets complex tasks requiring long context and extended output.
DeepSeek had previously notified developers of the price changes by email, with the new rates taking effect immediately.
V4 Pro also runs DSpark, an inference-acceleration framework jointly developed by DeepSeek and Peking University, which DeepSeek says improves generation speed by 57% to 78%.
From ‘cheap AI’ to commercial segmentation
The launch comes just a week after DeepSeek announced broader API price increases, signaling a shift away from its earlier ultra-low-cost strategy toward tiered monetization.
With leading AI models now competing across context windows ranging from 20K to 1M tokens, V4 Pro is betting that developers will increasingly evaluate models not simply by cost per token, but by useful tokens generated per unit of compute.


