gemini Gemini Blog ·

Gemini 3.8 Flash and Flash-Lite TTS Models Offer Advanced Voice Generation

aigaengineer
feature

Google has introduced two new text-to-speech models, Gemini 3.8 Flash TTS and 3.8 Flash-Lite TTS, designed to transform voice generation into a dynamic creative studio. These models allow creators, developers, and enterprises to create bespoke voices, customize roles, accents, and characteristics, and direct performance line by line with granular control. Key features include generative voice design, an expansive voice library, and voice replication with built-in consent verification and SynthID watermarking. The models are rolling out starting today for developers via the Gemini API and Google AI Studio, with enterprise availability coming soon.

  • →Introducing Gemini 3.8 Flash and Flash-Lite TTS Models
  • →Advanced Generative Voice Design and Customization
  • →Granular Control Over Performance and Delivery
  • →Built-in Trust, Consent, and Transparency Features
  • →Developer Access via AI Studio and Partner Integrations
Features (3) ›
  • Introducing Gemini 3.8 Flash and Flash-Lite TTS Models

    Google has launched two new text-to-speech models, Gemini 3.8 Flash TTS and 3.8 Flash-Lite TTS, which transform voice generation into a dynamic creative studio. These models enable richer, more expressive audio experiences for creators, developers, and enterprises, enhancing products like Gemini Notebook and Google Vids.

  • Advanced Generative Voice Design and Customization

    Gemini 3.8 Flash TTS allows users to create bespoke voices from scratch by customizing role, accent, and voice characteristics across over 100 languages using natural language prompts. It also offers access to 2,000+ production-ready voices and voice replication from just a 30-second audio sample, backed by consent verification.

  • Granular Control Over Performance and Delivery

    Both TTS models provide precise line-by-line control over how each line is delivered, enabling users to write stage directions or let Gemini steer delivery with natural script cues. Capabilities include long-form generation, native two-speaker scene staging, and scripted vocal bursts and backchanneling for realistic conversational texture.

Enhancements (1) ›
  • Built-in Trust, Consent, and Transparency Features

    For voice replication, the system leverages consent verification, requiring a verbal consent recording from the voice owner. Additionally, every audio clip generated by Gemini Audio models is watermarked with SynthID and C2PA credentials, ensuring AI-generated speech remains detectable to prevent misinformation.

Notes (2) ›
  • Developer Access via AI Studio and Partner Integrations

    Developers can try the new speech generation capabilities starting today in Google AI Studio's audio playground, which functions as a voice design workspace. The Gemini API supports integration with developer platforms like Agora, LiveKit, Pipecat, and Vercel, with partners such as Figma and HeyGen already leveraging the models.

  • General Availability for Developers, Enterprise Soon

    Gemini 3.8 Flash TTS and Flash-Lite TTS are rolling out starting today for developers via the Gemini API and Google AI Studio. Enterprise availability is coming soon via API in Gemini Enterprise, and the models are integrated into Gemini Notebook and Google Vids.

Read the original announcement →

https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-text-to-speech/

Related releases