iFlytek launches China’s first open-source million-token edge AI model

  • The 1.7B and 4B models bring long-context and agentic capabilities to phones, PCs, robots and other local devices
  • The launch marks a shift in edge AI from understanding commands to executing tasks, while making million-token context available through an open-source model

iFlytek (科大讯飞), one of China’s early AI pioneers, launched and open-sourced on September 1 two general-purpose edge AI models, Spark X2.5-4B and Spark X2.5-1.7B, through its wholly owned subsidiary Spark Yuan.

The Hefei-based company said the models are the first open-source edge models to natively support up to 1 million tokens of context, bringing a capability previously associated mainly with cloud-based models to devices such as smartphones, PCs and robots.

Benefits of edge AI

Edge AI refers to running AI models directly on local devices — rather than sending every request to the cloud.

That can improve privacy because data does not need to leave the device, reduce reliance on network connections and lower cloud-computing costs for companies serving large numbers of users.

The million-token context window also changes how smaller local models handle long documents.

Instead of splitting a lengthy technical manual into sections and processing them separately, Spark X2.5 can take in an entire document at once and maintain context across a conversation.

But the bigger shift is from understanding to execution.

Various applications

In office applications, the models can generate Chinese reports from raw data and translate them into English.

In smart-home applications, iFlytek says the models achieve a 90.3% command-execution accuracy with a response time of just 0.85 seconds.

The models can also run on robots to support operational control and navigation decisions, reducing their dependence on cloud-based AI.

In other words, long context addresses whether an AI can “see the whole picture,” while agentic capabilities and tool use determine whether it can do something with what it sees.

Image credit: iFlytek

The two models were trained entirely on domestic computing platforms using about 20 trillion tokens of training data, according to iFlytek.

The smaller 1.7B model, despite having a fraction of the parameters of many cloud-based models, can match models two to three times its size on tasks including code development, the company said.

The model weights are available for free on platforms including Hugging Face and GitHub, allowing developers to download and deploy them locally.

From cloud AI to devices

The launch comes as the AI industry moves toward putting increasingly capable models directly onto consumer electronics and machines.

Apple and Google have helped define the global edge-AI race, with Apple’s iPhone 16 marking a major step in bringing generative AI capabilities onto smartphones.

iFlytek is now betting that China can compete not only by developing larger cloud models, but by making smaller models capable enough to operate independently on everyday devices.

That distinction could become increasingly important as AI moves from the cloud into smartphones, cars, robots and industrial equipment.

Local models can avoid the latency, connectivity requirements and recurring inference costs associated with sending workloads to remote data centers.

Key to global expansion

The open-source approach also gives iFlytek a way to extend its technology beyond its own hardware and applications, allowing developers to deploy and adapt the models themselves.

iFlytek Chairman Liu Qingfeng (刘庆峰) has said that solving real-world “must-have” problems is a key competitive advantage for Chinese AI companies seeking to expand globally.

As AI competition moves from a race over cloud-based model size toward a contest over deployment, the companies that can make smaller models genuinely useful on local devices may gain an early advantage in the next generation of smart hardware.

Header image generated by Doubao