gemini Gemini Blog ·

Google introduces Gemini 3.5 Transcribe for precise, intelligent speech-to-text

aipreviewengineer
feature

Google has launched Gemini 3.5 Transcribe, its most precise speech-to-text model, designed to convert raw audio into accurate, polished, and formatted text. It significantly improves transcription accuracy, achieving a Word Error Rate (WER) of 2.6%–4.0%, and offers 70% faster latency compared to its predecessor, Chirp 3. Developers can access the model via the Gemini API for building voice agents, real-time captioning, and post-call analytics. The model is also integrated into Google products like the Gemini app and Gboard, and is available in public preview for developers and enterprises.

  • Gemini 3.5 Transcribe model introduced for intelligent speech-to-text
  • Available via Gemini API with real-time streaming and pre-recorded audio processing
  • Enhanced precision, smart transcription, and multi-language support
  • Integrated into Gemini app, Gboard, Antigravity, and upcoming Chrome features
  • Public preview for developers and enterprises now available
Features (3)
  • Gemini 3.5 Transcribe model introduced for intelligent speech-to-text

    Google has launched Gemini 3.5 Transcribe, its most precise speech-to-text model, designed to convert raw audio directly into accurate, polished, and formatted text. It aims to overcome limitations of conventional models by handling background noise, complex jargon, and disfluency cleanup.

  • Available via Gemini API with real-time streaming and pre-recorded audio processing

    Developers can now integrate Gemini 3.5 Transcribe via the Gemini API in Google AI Studio and Gemini Enterprise Agent Platform. It offers two APIs: a Live API for continuous, bidirectional streaming with sub-second latency, and an Interactions API for transcribing recorded audio with speaker attribution and word-level timestamps.

  • Enhanced precision, smart transcription, and multi-language support

    The model features smart transcription, which handles self-corrections and removes filler words, and achieves an average Word Error Rate (WER) of 2.6%–4.0%. It recognizes specialized jargon via custom vocabulary, automatically detects and transcribes over 85 languages, and includes multi-speaker identification for up to three speakers.

Notes (2)
  • Integrated into Gemini app, Gboard, Antigravity, and upcoming Chrome features

    Gemini 3.5 Transcribe enhances voice capabilities across Google products like the Gemini app on macOS, Gboard on Android with the new Rambler feature, and Google Antigravity. It is also slated for integration into Chrome, enabling talk-to-type functionality in web fields.

  • Public preview for developers and enterprises now available

    Gemini 3.5 Transcribe is available in public preview for developers in the Gemini API via Google AI Studio and Google Antigravity. It's also in public preview for enterprises via Gemini Enterprise Agent Platform, with consumer access through the Gemini app on macOS and Rambler on Android.

Read the original announcement →

https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5-transcribe/

Related releases