- Alibaba’s new model delivers near-flagship performance with just 6 billion active parameters
- It also sharply cuts training and inference costs, intensifying China’s AI price war
Alibaba has released and open-sourced Qwen3.8-Flash, a new-generation AI model that the company says delivers near-flagship performance while using a fraction of the computing resources required by its predecessor.
The model, unveiled Tuesday night, uses a new “Next” architecture with 125 billion total parameters but activates just 6 billion at a time. Alibaba says its performance surpasses Anthropic’s Claude Opus 4.6 and comes close to Opus 4.8.
The efficiency gains extend to both training and inference. Alibaba says Qwen3.8-Flash achieved performance comparable to Qwen3.7-Plus while requiring just one-ninth the computing resources to train, cutting training costs by 90%.
At inference, the model costs as little as 1 yuan ($0.14) per million input tokens and 3 yuan per million output tokens.
That puts its pricing at roughly 3% of Claude Opus 4.6 and, according to Alibaba, at two-thirds of DeepSeek-V4-Flash’s off-peak price and one-third of its peak price.
Optimized for agentic AI
Qwen3.8-Flash is optimized for agentic coding, professional office work and long-running tasks.
The Qwen and QwenWork teams have jointly optimized the model and its harness framework, further strengthening its ability to execute multi-step tasks.
The model is already integrated into QwenWork’s standard mode, where Alibaba says it can handle 95% of routine office tasks.
Technically, Qwen3.8-Flash is a multimodal mixture-of-experts model incorporating a GDN+QSA hybrid attention architecture, Gated Residual mechanisms and 51-billion-parameter N-gram embeddings.
Alibaba has also released Qwen3.8-Flash-Next, an early version of the architecture intended to give developers a head start on adaptation ahead of the next generation.
The model weights were released on Hugging Face and ModelScope at 11 p.m. on Aug. 26.
From bigger models to cheaper intelligence
Qwen3.8-Flash is the third major Qwen3.8 model to be open-sourced, following the 2.4-trillion-parameter Qwen3.8-Max and Qwen3.8-27B.
The broader Qwen family has surpassed 3 billion downloads globally, with more than 300,000 derivative models developed by the community, according to Alibaba.
The latest release highlights a broader shift in AI competition. As frontier models become increasingly capable, the key metric is moving beyond parameter counts and headline token prices toward the cost of completing an actual task.
By delivering near-frontier performance while activating only a small fraction of its parameters, Qwen3.8-Flash is betting that architectural efficiency — rather than sheer model size — will determine the next phase of the AI cost race.
For international developers, the combination of open weights, low inference costs and agentic capabilities could make Qwen3.8-Flash a more practical alternative for building AI applications at scale.
Header image credit: Alibaba Qwen


