SevenTnewS

Speech synthesis

Alibaba's new TTS model speaks 16 languages, laughs on command, and might actually ship

Alibaba's Qwen-Audio-3.0-TTS is a production-oriented speech synthesis system that combines a low-frame-rate tokenizer with progressive training. It supports 16 languages, 20 Chinese dialect regions, and natural-language instructions for emotion, pace, and speaking style.

Emmanuel Fabrice Omgbwa Yasse Assisté par IA

2026-07-26 · 4 min de lecture

Alibaba's new TTS model speaks 16 languages, laughs on command, and might actually ship
Sources : Qwen-Audio-3.0-…

Alibaba's speech synthesis team released Qwen-Audio-3.0-TTS, a system that tries to solve nearly every problem that keeps text-to-speech models out of production: latency, unnatural prosody, limited language support, and the inability to follow anything but the most rigid instructions. The paper, published by the Funaudiollm group at Alibaba Group, claims first place on the independent Artificial Analysis Text-to-Speech Leaderboard and state-of-the-art results across a broad set of evaluations. This aligns with Alibaba's broader push to integrate AI into its $53 billion infrastructure bet, per their latest update.

Low-frame-rate tokenizer shrinks latency

The core technical bet is a 12.5 Hz supervised speech tokenizer. Most speech systems operate at much higher frame rates, which means autoregressive decoding takes longer. Running at 12.5 Hz cuts that cost while, according to the paper, still preserving enough content and speaker information. The model combines this tokenizer with a five-stage progressive training pipeline that separately optimizes the language model and the foundational model before joining them for reinforcement learning.

This is not the first time Alibaba has pushed the boundaries of audio generation. Earlier this year, the company released Qwen-Music, which composes melodies and fills in the band arrangement before generating vocals. Qwen-Audio-3.0-TTS focuses on the voice side only, but it overlaps with the company's broader push toward multimodal systems, including Qwen2.5-Omni, which handles real-time speech and text interaction.

Natural-language controls and fine-grained tags

The most visible feature is controllability. The model accepts free-style natural-language instructions: "say this slowly in a calm voice," "read the following with disgust," or "speak quickly like you are in a hurry." On top of that, the paper introduces 86 inline tags for phrase- and word-level control. These cover expressive transitions (sarcastic, panicked, excited) and non-verbal events such as laughter, breathing, coughing, and sighing. The samples in the paper demonstrate a single sentence shifting from a flat read to a giggly, mischievous tone or an angry outburst, depending on the tag inserted.

This degree of granularity is unusual. Most TTS models offer at best a handful of preset voices or emotions. The inline tags effectively let writers or sound designers annotate a script like a director marking up a script for an actor. Such controllability echoes the philosophy behind effective prompt engineering in general, where precise instructions can dramatically improve output quality.

Language and dialect coverage

The model supports 16 languages, seven of them new to the Qwen-Audio line. The additions include Arabic, Indonesian, Portuguese, Thai, Vietnamese, Malay, and Tagalog. For Chinese, Qwen-Audio-3.0-TTS covers 20 dialect regions, from Shanghainese to Cantonese to Northeastern Mandarin. The paper provides side-by-side samples showing zero-shot voice cloning across all 14 languages the CosyVoice 3.0 baseline also supports, plus the seven new ones.

Cross-lingual synthesis works too: a prompt in Mandarin can produce speech in English, Japanese, or Korean while preserving the original speaker's timbre. The evaluation tables show that the model maintains voice consistency across languages better than the CosyVoice baselines in most tested conditions. This kind of cross-lingual ability is becoming a hallmark of Alibaba's approach, as seen in their robotics models that also span multiple domains.

Robustness and long-form synthesis

One of the paper's more pragmatic contributions is acoustic robustness. The model can synthesize clean speech from noisy, reverberant, or telephone-bandwidth reference audio. The paper shows examples where the source recording has background noise, echo, or muffled frequency cuts, and the output remains intelligible and natural. There is no explicit denoising mode; the model apparently learns to ignore degradation during training.

Long-form synthesis works in a single pass for up to three minutes. The provided samples read multi-paragraph English and Chinese text without the wobbling quality or dropped prosody that shorter-segment approaches often produce. The model handles hard text-normalization cases too: numbers, abbreviations, symbols, and named entities all receive correct readings in the examples.

Verdict: 8/10

  • Real price: Not yet publicly priced. The paper describes a production-oriented system from Alibaba, but no API tier or cost is disclosed as of publication.
  • Ideal for: Teams building voice assistants or dubbing pipelines that need fine-grained emotional control and broad language coverage.
  • Avoid if: You need offline inference or a self-hostable open-weight model. The paper mentions a reproducible speaker fine-tuning protocol and vocoder super-resolution, but the full model weights have not been released.
  • Alternatives: CosyVoice 3.0 for an already-deployed Chinese-dominant system with comparable quality; Nvidia Audex if you need a single model that also handles speech recognition and audio generation (see Nvidia's Audex).
  • Test date: The paper's evaluations are based on internal tests conducted before the March 2025 submission date. Independent third-party reproduction has not yet been published.

The paper is thorough and its benchmarks are convincing. But the usual caveats apply: the samples are cherry-picked, the human evaluation arena uses Alibaba's own test sets, and the one-shot voice cloning demos for the newly supported languages lack any comparison because the CosyVoice baseline does not support those languages at all. The real test will be how the model performs when deployed at scale against real-world noise and accent variation. Until the weights or an API are public, the leaderboard ranking is a useful signal but not a final verdict.

L'essentiel de la tech en 3 minutes chaque matin

Un email, chaque jour ouvré, avec ce qui compte vraiment en IA et en tech.