Skip to main content
LLMgram · AI News · 2026-08-26

Google launches Gemini 3.5 Transcribe with 85-language STT and realtime streaming

Google launches Gemini 3.5 Transcribe with 85-language STT and realtime streaming

Google has launched Gemini 3.5 Transcribe, expanding its Gemini model family into speech-to-text with capabilities aimed at production workloads rather than basic dictation. The model supports smart transcription, function calling, custom vocabulary, multi-speaker identification, and more than 85 languages with automatic language detection out of the box. Google also positions it as more accurate than prior options, citing lower word error rate, and adds realtime streaming for live audio pipelines. Initial availability spans Google AI Studio, developer APIs, and Gemini Desktop on macOS, giving builders an immediate path to test and integrate voice workflows. Pricing tiers, latency benchmarks, and head-to-head accuracy against rival STT services remain unconfirmed in the initial announcement.

Sources

Google launches Gemini 3.5 Transcribe with 85-language STT and realtime streaming

Google launches Gemini 3.5 Transcribe with 85-language STT and realtime streaming

Google announced Gemini 3.5 Transcribe, a new speech-to-text model with smart transcription, function calling, lower WER, custom vocabulary, multi-speaker identification, and support for over 85 languages. The release also includes realtime streaming support.

Key takeaway

Gemini 3.5 Transcribe adds agent-ready STT with function calling, streaming, and 85+ languages to Google AI Studio and APIs now.

What happened

Google announced Gemini 3.5 Transcribe, a new speech-to-text model with smart transcription, function calling, lower word error rate, custom vocabulary, multi-speaker identification, and support for over 85 languages.

Reporting tied to the launch states the model is available on Google AI Studio and APIs as well as Gemini Desktop for macOS, and includes realtime streaming support with auto-detection of more than 85 languages out of the box.

Evidence

  • Google announced Gemini 3.5 Transcribe as a new speech-to-text model with smart transcription, function calling, lower WER, custom vocabulary, multi-speaker identification, and 85+ language support.

    X · attributed

    Google announced Gemini 3.5 Transcribe, a new speech-to-text model with smart transcription, function calling, lower WER, custom vocabulary, multi-speaker identification, and support for over 85 languages.

  • The release includes realtime streaming support.

    X · attributed

    The release also includes realtime streaming support.

  • Gemini 3.5 Transcribe is available on Google AI Studio, APIs, and Gemini Desktop for macOS.

    X · attributed

    A new Gemini 3.5 Transcribe is now available on Google AI Studio and APIs as well as Gemini Desktop for macOS!

  • The model offers auto-detection of 85+ languages out of the box.

    X · attributed

    Auto-detection of 85+ languages out of the box

Why it matters

By embedding function calling and streaming inside transcription, Google is positioning STT as an agent interface layer, not just a text dump from audio.

Limits and uncertainties

The packet provides no pricing, latency figures, or independent benchmark comparisons for the lower WER claim.

Availability details beyond Google AI Studio, APIs, and Gemini Desktop for macOS are not specified in the evidence.

Practical implications

Teams building voice agents can evaluate Gemini 3.5 Transcribe immediately through Google AI Studio or APIs instead of waiting for a separate documentation drop.

Apps needing multilingual input can prototype with automatic detection across 85+ languages, but should validate accuracy on their target locales before production rollout.

What to watch

Whether Google publishes formal API documentation, pricing, and benchmark data beyond the initial X announcement.

Comparative WER and streaming latency results from third-party testers against Whisper, Deepgram, and other STT providers.

Sources

LLMgram editorial selection and synthesis · @llmgram. LLMgram is not the original publisher of this information.
Continue on LLMgram: Open in AI Signal →
Original reporting: Introducing Gemini 3.5 Transcribe, our new speech to text model with smart transcription, function calling, more precise transcription (lower WER), custom vocabulary support, multi