Google launches Gemini 3.5 Transcribe with 85-language STT and realtime streaming
Google has launched Gemini 3.5 Transcribe, expanding its Gemini model family into speech-to-text with capabilities aimed at production workloads rather than basic dictation. The model supports smart transcription, function calling, custom vocabulary, multi-speaker identification, and more than 85 languages with automatic language detection out of the box. Google also positions it as more accurate than prior options, citing lower word error rate, and adds realtime streaming for live audio pipelines. Initial availability spans Google AI Studio, developer APIs, and Gemini Desktop on macOS, giving builders an immediate path to test and integrate voice workflows. Pricing tiers, latency benchmarks, and head-to-head accuracy against rival STT services remain unconfirmed in the initial announcement.
Google launches Gemini 3.5 Transcribe with 85-language STT and realtime streaming
Google announced Gemini 3.5 Transcribe, a new speech-to-text model with smart transcription, function calling, lower WER, custom vocabulary, multi-speaker identification, and support for over 85 languages. The release also includes realtime streaming support.
Key takeaway
Gemini 3.5 Transcribe adds agent-ready STT with function calling, streaming, and 85+ languages to Google AI Studio and APIs now.
What happened
Google announced Gemini 3.5 Transcribe, a new speech-to-text model with smart transcription, function calling, lower word error rate, custom vocabulary, multi-speaker identification, and support for over 85 languages.
Reporting tied to the launch states the model is available on Google AI Studio and APIs as well as Gemini Desktop for macOS, and includes realtime streaming support with auto-detection of more than 85 languages out of the box.
Evidence
Google announced Gemini 3.5 Transcribe as a new speech-to-text model with smart transcription, function calling, lower WER, custom vocabulary, multi-speaker identification, and 85+ language support.
X · attributed
Google announced Gemini 3.5 Transcribe, a new speech-to-text model with smart transcription, function calling, lower WER, custom vocabulary, multi-speaker identification, and support for over 85 languages.
The release includes realtime streaming support.
X · attributed
The release also includes realtime streaming support.
Gemini 3.5 Transcribe is available on Google AI Studio, APIs, and Gemini Desktop for macOS.
X · attributed
A new Gemini 3.5 Transcribe is now available on Google AI Studio and APIs as well as Gemini Desktop for macOS!
The model offers auto-detection of 85+ languages out of the box.
X · attributed
Auto-detection of 85+ languages out of the box
Why it matters
By embedding function calling and streaming inside transcription, Google is positioning STT as an agent interface layer, not just a text dump from audio.
Limits and uncertainties
The packet provides no pricing, latency figures, or independent benchmark comparisons for the lower WER claim.
Availability details beyond Google AI Studio, APIs, and Gemini Desktop for macOS are not specified in the evidence.
Practical implications
Teams building voice agents can evaluate Gemini 3.5 Transcribe immediately through Google AI Studio or APIs instead of waiting for a separate documentation drop.
Apps needing multilingual input can prototype with automatic detection across 85+ languages, but should validate accuracy on their target locales before production rollout.
What to watch
Whether Google publishes formal API documentation, pricing, and benchmark data beyond the initial X announcement.
Comparative WER and streaming latency results from third-party testers against Whisper, Deepgram, and other STT providers.
Original reporting: Introducing Gemini 3.5 Transcribe, our new speech to text model with smart transcription, function calling, more precise transcription (lower WER), custom vocabulary support, multi