gemini Gemini Blog ·

Google introduces Gemini 3.8 Live models for advanced real-time voice AI

aigaengineer
feature

Google has introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two new models designed to enhance near real-time reasoning for voice agents. These models aim to make AI conversations more intuitive, intelligent, and fluid by combining conversational intelligence with visual grounding and multi-step reasoning. They are available today for developers via the Gemini API and Google AI Studio, with private previews for enterprises and broader rollout for Google Workspace and Search users. The release empowers developers and enterprises to build production-ready voice-driven interfaces.

  • Introduce Gemini 3.8 Live models for enhanced voice agents
  • New capabilities: visual input, multilingual detection, and background tool execution
  • Simultaneous reasoning and live progress narration for complex tasks
  • Achieve leading performance in speech-to-speech quality and task completion
  • Empowering the developer and enterprise voice ecosystem
Features (3)
  • Introduce Gemini 3.8 Live models for enhanced voice agents

    Google announces Gemini 3.8 Live and 3.8 Live Extended Thinking, new models designed to bring advancements in near real-time reasoning for more intuitive and intelligent AI voice agents. Gemini 3.8 Live is built for scale and cost efficiency with conversational intelligence and visual grounding, while Extended Thinking handles high-complexity tasks with increased intelligence and multi-step reasoning.

  • New capabilities: visual input, multilingual detection, and background tool execution

    Gemini 3.8 Live processes visual inputs in near real-time, enriching conversations with context. It automatically detects and transitions between 97 supported languages mid-conversation and executes tools/API calls in the background while maintaining dialogue flow.

  • Simultaneous reasoning and live progress narration for complex tasks

    For deeper reasoning tasks, 3.8 Live Extended Thinking reasons and speaks simultaneously, providing increased intelligence for complex workflows without interrupting conversation. It uses early verbal cues and live progress narration to guide users through multi-step background tasks.

Enhancements (2)
  • Achieve leading performance in speech-to-speech quality and task completion

    Gemini 3.8 Live Extended Thinking achieved the #1 spot on Artificial Analysis' Speech to Speech Quality Index and leads in agentic task completion on τ-Voice and Sierra’s τ-Voice-banking benchmark. Gemini 3.8 Live secured second place in the Speech Agent Arena, demonstrating high user preference and cost-effectiveness.

  • Ensure transparency with SynthID audio watermarking

    All audio generated by Google's AI products, including the new Gemini 3.8 Live models, is imperceptibly watermarked with SynthID. This feature helps ensure AI-generated content remains detectable to prevent misinformation.

Notes (2)
  • Empowering the developer and enterprise voice ecosystem

    The Gemini Live API enables developer platforms like Agora, LiveKit, and Vercel to build and deploy high-performance voice-driven interfaces. Google is also partnering with companies such as Salesforce, Genspark, and Lumeris, who highlight the models' impressive latency, fluidity, and tool-calling capabilities.

  • Immediate availability for developers and phased rollout for enterprises and users

    Gemini 3.8 Live and 3.8 Live Extended Thinking are rolling out today for developers via the Gemini API and Google AI Studio. Enterprises can access private previews, with broader availability coming soon to Gemini Enterprise, Google Workspace, and for all Google AI Pro and Ultra subscribers in various Google products.

Read the original announcement →

https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking/

Related releases