- The new 552-billion-parameter model offers native multimodal vision while cutting memory and hardware requirements
- Tencent’s WorkBuddy has integrated the model, joining CodeBuddy and OpenCode in offering access
DeepSeek on September 10 launched V4.1 Flash, the smallest model in its new architecture series, with native multimodal vision capabilities and a focus on faster inference, higher throughput and lower operating costs.
The model uses a 552-billion-parameter mixture-of-experts architecture and has undergone a new pretraining approach and more extensive reinforcement-learning post-training.
DeepSeek said it outperformed several flagship models, including DeepSeek V4 Pro, GLM-5.3 and Kimi K3, in benchmark tests, while generating responses nearly twice as fast.

Its architecture also significantly reduces memory requirements. DeepSeek said V4.1 Flash cuts demand for high-bandwidth memory to one-quarter of the previous generation and SSD storage to one-eighth, potentially lowering the cost of running agent-based AI tasks.
DeepSeek has consequently cut API prices for V4.1 Flash while retaining its peak/off-peak pricing model, with off-peak rates set at half the peak price. The new rates took effect at noon on September 10.
WorkBuddy integration
As an official DeepSeek partner, Tencent’s WorkBuddy has fully integrated V4.1 Flash in China and is offering a two-week promotional discount.
CodeBuddy and OpenCode have also added support.
Users can access the model by changing the model name to deepseek-flash. The previous V4 Flash and V4 Flash Vision Exp models have been discontinued, with their existing model names temporarily routed to V4.1 Flash for compatibility.
A single entry point
As part of the update, DeepSeek’s web version has also consolidated its previous “Fast,” “Expert” and “Vision” modes into a single entry point.
Users no longer need to choose a mode manually: the AI automatically selects the appropriate mode based on task complexity and activates its vision capabilities when an image is detected.

