DeepSeek cuts Flash API prices by up to 60% as AI price war heats up

  • The new pricing takes effect September 10, with the biggest cut for cached inputs during off-peak hours
  • The move comes a day after DeepSeek began testing a faster, native-multimodal V4.1 Flash model that could challenge its more powerful Pro version

DeepSeek is cutting prices for its V4 Flash application programming interface (API) by as much as 60%, extending its push to make AI inference cheaper while using time-based pricing to manage demand for computing power.

The new rates will take effect at noon Beijing time on September 10. During off-peak hours, the price for cached input will fall from 0.05 yuan to 0.02 yuan per million tokens, while uncached input will drop from 1.5 yuan to 1 yuan. Output prices will fall from 4.5 yuan to 4 yuan.

Peak-hour rates for cached and uncached input and output remain 0.1 yuan, 3 yuan and 9 yuan.

Cached input refers to text that the model has already processed and can reuse, making it much cheaper to serve. New input requires fresh computation and therefore costs more.

Image credit: DeepSeek

Greater use of caching

DeepSeek will charge twice as much during peak periods — 9 a.m. to noon and 2 p.m. to 6 p.m. on weekdays — effectively encouraging developers to run less time-sensitive workloads during evenings and weekends.

The pricing structure also puts a particularly strong incentive on developers to design applications that make greater use of caching, which can reduce the amount of computing needed for repeated requests.

For developers, the effective reduction in total API costs could be around 40%, although the actual savings will depend on factors including cache-hit rates.

Flash gets faster

The price cut comes just a day after DeepSeek opened internal testing of V4.1 Flash, an intermediate version with a new architecture and native multimodal capabilities.

Developers testing the model have reported peak output speeds of 507 tokens per second, with average speeds exceeding 300 tokens per second in multiple tests.

DeepSeek has described the new model as more capable, faster and cheaper, and its internal feedback survey asks users whether it could fully replace the current V4 Pro.

That suggests DeepSeek may be trying to push Flash beyond its traditional role as a lightweight, low-cost model and closer to Pro-level capabilities.

Tiered pricing strategy

The latest cut also reverses a price hike announced less than a month ago, when the launch of V4 Pro sent some cached-input prices as much as 11 times higher and triggered complaints from developers.

DeepSeek’s latest strategy combines aggressive pricing to win users and peak/off-peak rates to encourage better use of limited compute.

For global AI developers, the pressure is significant. DeepSeek’s V4 Flash scores about 53 on Artificial Analysis’ intelligence index, below GPT-5.6 Sol at 61 and Claude Opus 5 at 63.

But its estimated cost per task is far lower — about $0.25 versus $1.23 and $2.34 respectively.

Aggressive pricing to win customers

The gap highlights DeepSeek’s core proposition: even if its models do not lead on raw capability, substantially lower inference costs can make them attractive for developers running AI at scale.

For consumers, the impact will be less immediate. Lower API costs could eventually translate into cheaper AI services or more generous usage limits, but any savings would first have to pass through developers and application providers.

Header image generated by Tencent Yuanbao