LLMgram · AI News · 2026-07-23

Qwen launches Audio-3.0-TTS with Flash and Plus modes

Qwen launches Audio-3.0-TTS with Flash and Plus modes

Alibaba’s Qwen team introduced Qwen-Audio-3.0-TTS, a text-to-speech model offered in Flash for real-time interaction and Plus for higher-quality generation. The release adds fine-grained inline tags for emotional and physical performance cues in generated speech.

Key takeaway

Qwen is treating voice as a first-class product surface with latency-versus-quality SKUs, not a single generic TTS endpoint.

Context

The Flash and Plus split targets builders who need low-latency conversational agents versus offline or premium narration quality from the same model family. That packaging mirrors how major labs now ship modality models as application-ready variants rather than research demos.

Inline emotion and physical-cue tags reduce the need for separate prosody controllers or multi-model pipelines when injecting expressive speech into agents and media workflows. Relative to earlier Qwen image releases, this extends the Studio stack into audio generation with controllable delivery.

LLMgram editorial selection and synthesis · @llmgram. LLMgram is not the original publisher of this information.
Continue on LLMgram: Open in AI Signal →
Original reporting: Introducing the Qwen-Audio-3.0-TTS. Our latest text-to-speech model, in two flavors: • Flash: real-time interaction • Plus: high-quality generation What's new: • Fine-grained inli…