Google introduces Gemini 3.5 Transcribe for precise, intelligent speech-to-text
Google has launched Gemini 3.5 Transcribe, its most precise speech-to-text model, designed to convert raw audio into accurate, polished, and formatted text. It significantly improves transcription accuracy, achieving a Word Error Rate (WER) of 2.6%–4.0%, and offers 70% faster latency compared to its predecessor, Chirp 3. Developers can access the model via the Gemini API for building voice agents, real-time captioning, and post-call analytics. The model is also integrated into Google products like the Gemini app and Gboard, and is available in public preview for developers and enterprises.
- →Gemini 3.5 Transcribe model introduced for intelligent speech-to-text
- →Available via Gemini API with real-time streaming and pre-recorded audio processing
- →Enhanced precision, smart transcription, and multi-language support
- →Integrated into Gemini app, Gboard, Antigravity, and upcoming Chrome features
- →Public preview for developers and enterprises now available
Features (3) ›
- Gemini 3.5 Transcribe model introduced for intelligent speech-to-text
Google has launched Gemini 3.5 Transcribe, its most precise speech-to-text model, designed to convert raw audio directly into accurate, polished, and formatted text. It aims to overcome limitations of conventional models by handling background noise, complex jargon, and disfluency cleanup.
- Available via Gemini API with real-time streaming and pre-recorded audio processing
Developers can now integrate Gemini 3.5 Transcribe via the Gemini API in Google AI Studio and Gemini Enterprise Agent Platform. It offers two APIs: a Live API for continuous, bidirectional streaming with sub-second latency, and an Interactions API for transcribing recorded audio with speaker attribution and word-level timestamps.
- Enhanced precision, smart transcription, and multi-language support
The model features smart transcription, which handles self-corrections and removes filler words, and achieves an average Word Error Rate (WER) of 2.6%–4.0%. It recognizes specialized jargon via custom vocabulary, automatically detects and transcribes over 85 languages, and includes multi-speaker identification for up to three speakers.
Notes (2) ›
- Integrated into Gemini app, Gboard, Antigravity, and upcoming Chrome features
Gemini 3.5 Transcribe enhances voice capabilities across Google products like the Gemini app on macOS, Gboard on Android with the new Rambler feature, and Google Antigravity. It is also slated for integration into Chrome, enabling talk-to-type functionality in web fields.
- Public preview for developers and enterprises now available
Gemini 3.5 Transcribe is available in public preview for developers in the Gemini API via Google AI Studio and Google Antigravity. It's also in public preview for enterprises via Gemini Enterprise Agent Platform, with consumer access through the Gemini app on macOS and Rambler on Android.
https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5-transcribe/
Related releases
- Google showcases 7 ways Gemini enhances student productivity in Workspace Gemini Blog ·
- Gemini Live adds new voice-activated productivity features with Spark integration Gemini Blog ·
- Google's Python GenAI SDK v2.20.0 adds audio transcription mode and JPEG2000 support Gemini Python SDK Releases ·
- Genkit CLI Developer UI enhanced with improved logging and error visibility Genkit Releases ·
- Gemini for macOS: How to Use Intelligent Dictation Gemini Blog ·
- Gemini model gemini-robotics-er-1.6-preview reaches end of life in 7 days endoflife.date ·