Qwen launches Audio-3.0-TTS with Flash and Plus modes

Alibaba’s Qwen team introduced Qwen-Audio-3.0-TTS, a text-to-speech model offered in Flash for real-time interaction and Plus for higher-quality generation. The release adds fine-grained inline tags for emotional and physical performance cues in generated speech.
Key takeaway
Qwen is treating voice as a first-class product surface with latency-versus-quality SKUs, not a single generic TTS endpoint.
Context
The Flash and Plus split targets builders who need low-latency conversational agents versus offline or premium narration quality from the same model family. That packaging mirrors how major labs now ship modality models as application-ready variants rather than research demos.
Inline emotion and physical-cue tags reduce the need for separate prosody controllers or multi-model pipelines when injecting expressive speech into agents and media workflows. Relative to earlier Qwen image releases, this extends the Studio stack into audio generation with controllable delivery.