DeepSeek begins testing V4.1 Flash with native multimodal support

  • The new model uses a redesigned architecture that DeepSeek says is faster, cheaper and more capable
  • The limited test also hints at an ambition to bring Flash-level speed and pricing closer to the capabilities of its Pro models

DeepSeek has quietly begun internal testing of DeepSeek V4.1 Flash, an intermediate release that introduces a new model architecture and native multimodal capabilities.

The Chinese AI upstart opened the test to developers through its official community groups on September 8. Users can access the model without changing their API endpoint, simply by switching the model name to “deepseek-v4.1-flash-expires-on-0910.”

Pricing remains the same as the existing V4 Flash.

The endpoint is clearly designed for testing rather than production use. It expires on September 10 and supports a maximum of 20 concurrent requests per account, compared with 2,500 for the production version.

A new architecture

The most notable change is native multimodal support. Previously, DeepSeek required a separate vision extension for mixed text-and-image inputs, with the V4-Flash-Vision-Exp model released August 21 as an experimental solution.

Native multimodality means the capability is built into the model itself, allowing it to handle text and images without an additional vision module or switching between models.

The new architecture also marks a departure from the V4-Flash-0731 released in late July, which focused on post-training improvements without changing the underlying architecture.

DeepSeek has not disclosed what specifically changed in V4.1 Flash.

Image credit: Screenshot from DeepSeek user interface

More revealingly, the company’s internal testing questionnaire asks users whether the model could fully replace the current DeepSeek V4 Pro online — suggesting the goal may be more ambitious than simply making Flash faster.

Another cost-killer?

DeepSeek has made low-cost inference a central part of its challenge to global AI incumbents.

In June, the company said V4.1 would cut prices by another 15%. At the time, V4 Flash cost $0.14 per million input tokens and $0.28 per million output tokens, compared with $5 and $25 for Anthropic’s Claude Opus 4.8, according to the company.

If V4.1 Flash can narrow the performance gap with flagship models while retaining Flash-level pricing, it could put further pressure on the pricing strategies of OpenAI, Anthropic and other leading AI providers.

For now, however, the test endpoint expires within days, and its actual performance and pricing will only become clear if DeepSeek turns the experiment into a full release.

Header image credit: DeepSeek