Qwen-Audio-3.0-TTS
Qwen-Audio-3.0-TTS
Qwen Audio, Token Foundry, Alibaba Group
funaudiollm@alibaba-inc.com
Abstract:<br>In this paper, we present Qwen-Audio-3.0-TTS, a production-oriented speech synthesis system that jointly advances content consistency, speaker similarity, prosodic naturalness, audio quality, controllability, multilingual coverage, efficiency, and robustness. It combines a 12.5 Hz low-frame-rate speech tokenizer for reduced inference latency with a five-stage progressive training paradigm for coordinated LM and FM optimization. The model provides production-level control through free-style natural-language instructions and fine-grained inline tags, while supporting 16 languages, 20 Chinese dialect regions, one-pass long-form synthesis up to 3 minutes, and robust generation from noisy, reverberant, or unclear reference speech. Across SEED-TTS-Eval, CV3-Eval, instruction-following, long-form, and acoustic-robustness evaluations, Qwen-Audio-3.0-TTS achieves state-of-the-art performance on many reported dimensions or the strongest aggregate results. It also ranks first on the independent Artificial Analysis Text-to-Speech Leaderboard. These results establish Qwen-Audio-3.0-TTS as a strong foundation for production-level speech synthesis.
Key Contributions:
Low-frame-rate speech tokenizer: A 12.5 Hz supervised speech tokenizer reduces autoregressive decoding cost while retaining content and speaker information.
Progressive training paradigm: The training pipeline combines independent LM and FM pretraining, joint training with high-quality data annealing, LM reinforcement learning, FM robustness training, and FM reinforcement learning to improve content consistency, prosodic naturalness, voice fidelity, perceptual quality, and robustness.
Production-grade controllability: The model interprets free-style natural-language instructions describing role, emotion, speaking style, rate, timbre, and accent. In parallel, 86 newly added fine-grained inline tags enable localized control at phrase and word level, including expressive transitions and non-verbal events such as laughter, breathing, coughing, and sighing.
Broad and robust deployment coverage: The model supports 16 languages, seven of them newly added, and 20 Chinese dialect regions; it handles hard text-normalization cases, one-pass synthesis up to 3 minutes, and degraded prompts without an explicit denoising mode. A reproducible speaker fine-tuning protocol and vocoder super-resolution further support target-voice adaptation and 48 kHz output.
Comprehensive evaluation: We evaluate zero-shot voice cloning, multilingual and cross-lingual synthesis, free-style instruction following, fine-grained control, text normalization, long-form generation, adverse-prompt robustness, and dialect synthesis through objective benchmarks and arena-style human evaluation.
Content Consistency
Speaker Similarity
Contents
Zero-shot In-context Generation
Multilingual / Minority Languages
Cross-lingual In-context Generation
Emotional Voice Generation
Chinese Dialect Voice Generation
Acoustic Robustness
Long-text Generation
Text Normalization
Instructed Voice Generation
Fine-grained Tags
Zero-shot In-context Generation
Language<br>Prompt<br>Text<br>CosyVoice 3.0-0.5B<br>CosyVoice 3.0-1.5B<br>Qwen-Audio-3.0-TTS
ZH某些此类应用程序甚至让用户把智能手机对着现实世界中的标志或其他物体,就能翻译上面的外语文字。<br>ENSince then, the Brazilian has featured in 53 matches for the club in all competitions and has scored 24 goals.<br>HARD-ZH板凳宽,扁担长,板凳比扁担宽,扁担比板凳长,扁担要绑在板凳上,板凳不让扁担绑在板凳上,扁担偏要板凳让扁担绑在板凳上。<br>HARD-ENOh, no! I left my keys—again?! What am I going to do...call a locksmith?
Multilingual / Minority Languages
Zero-shot voice cloning across 14 languages.
Language<br>Prompt<br>Text<br>CosyVoice 3.0-0.5B<br>CosyVoice 3.0-1.5B<br>Qwen-Audio-3.0-TTS
JAこの種の人は論理的思考力に富み、パターンを記憶したり、問題を解決したり、科学的なテストに取り組んだりできます。<br>KO학교는 교육과정을 자의적으로 운영하거나 학생에게 임의적인 교내외 행사 참석을 강요하여서는 아니 된다.<br>DEDie häufigste Unfallursache im Winter sind rutschige Straßen, Gehwege (Bürgersteige) und vor allem Stufen.<br>ESAdemás pueden encontrarse varias especies más de robles de porte arbustivo en esta región.<br>FRDe nos jours, les ordinateurs sont utilisés pour modifier des photos et des vidéos.<br>ITÈ stata cancellata dopo la prima stagione.<br>RUПосещение этого места можно удобно совместить с лодочной прогулкой по озеру.<br>— The languages below are newly supported in Qwen-Audio-3.0-TTS and are not supported by CosyVoice 3.0. —<br>ARالوحدة المختارة تمنح فرصة للتأمل واستعادة التوازن الداخلي المفقود.——<br>IDPeriksa kesehatan rutin setahun sekali deteksi dini penyakit kronis sebelum parah.——<br>PTNossa série sobre artesanato tradicional apresenta hoje os mestres ceramistas do Vale do Jequitinhonha, que transformam o barro em verdadeiras obras de arte.——<br>THความเคารพผู้ใหญ่คือมารยาทไทยควรสืบทอด——<br>VIThiên nhiên chữa lành tuyệt vời nếu ta chịu bước chậm và lắng nghe tiếng gió.——<br>MSRancangan ini pada asalnya menampilkan pelakon suara amatur tempatan Texas Timur.——<br>TLNgayon, ang tanging mga insekto na hindi...