Qwen-Audio-3.1 tops global voice leaderboard as prices plunge

  • Qwen-Audio-3.1-TTS scores first on Artificial Analysis’ voice leaderboard
  • Alibaba cuts voice-model prices by up to 95% as it targets broader adoption

Alibaba’s latest text-to-speech model has taken the top spot on a global AI voice leaderboard, highlighting the company’s push to compete with leading overseas providers in speech generation.

Qwen-Audio-3.1-TTS scored 1,177 Elo points to rank first on Artificial Analysis’ Controlled Voice Arena on September 28.

The blind-preference benchmark evaluates models based on users’ preferences without revealing their identities to evaluators, according to Alibaba.

The model is part of Alibaba’s Qwen-Audio-3.1 series, unveiled at the 2026 Apsara Conference held in Hangzhou.

Matching delivery with context

The series also includes speech recognition and real-time voice interaction models, all of which are available through Qwen AI platform.

Unlike conventional text-to-speech systems that focus mainly on accurate pronunciation, Qwen-Audio-3.1-TTS is designed to match delivery with context.

Users can control factors such as emotion, speaking speed and expression through prompts, allowing the same voice to speak naturally in multiple languages and dialects.

The result is intended to move voice synthesis beyond simply “saying it correctly” toward “sounding right” for a particular context, Alibaba said in a notice.

Global implications

The ranking puts Alibaba ahead of established voice-AI companies including Cartesia, ElevenLabs and OpenAI on the benchmark, according to Artificial Analysis’ leaderboard.

Alibaba is also using price cuts to lower the barrier to adoption. It has reduced prices across its Qwen-Audio voice models by as much as 95%, including a roughly 70% cut for text-to-speech services.

The company’s open-source voice models, including Qwen-Audio-Agent and CosyVoice, have also attracted tens of thousands of stars on GitHub, reflecting interest from developers in China and overseas, Alibaba said.

For global developers, the combination of benchmark performance and lower API prices could make advanced voice generation cheaper to deploy at scale, as Chinese AI companies increasingly compete in specialized model markets.

Header image credit: Alibaba Qwen