gemini Gemini Blog ·

Gemini 3.8 Live Introduces Real-Time Live Avatar for Enterprise

aigaarchitect
feature

Google has launched Live Avatar for Gemini 3.8 Live, bringing near real-time visual presence to its native live dialogue models. This feature pairs real-time video generation with speech, enabling dynamic visual personas for interactive virtual offerings. It enhances customer service and walkthroughs by supporting precise lip-syncing, natural expressions, and fluid turn-taking. Live Avatar is now available in Gemini Enterprise, offering multimodal conversations, asynchronous tool execution, multilingual support across 97 languages, and custom avatar creation.

  • →Gemini 3.8 Live Introduces Real-Time Live Avatar
  • →Supports Asynchronous Tool Calls with Continuous Dialogue
  • →Allows Customization to Match Brand Needs
  • →Enables More Natural, Multimodal Interactions
  • →Offers Seamless Multilingual Support Across 97 Languages
Features (3) ›
  • Gemini 3.8 Live Introduces Real-Time Live Avatar

    The Live Avatar feature pairs near real-time video generation with speech, creating a dynamic visual persona that listens, sees, and speaks. It brings precise lip-syncing, natural expressions, and fluid turn-taking to enterprise virtual offerings, transforming digital exchanges into richer experiences.

  • Supports Asynchronous Tool Calls with Continuous Dialogue

    Live Avatar leverages Gemini’s advanced reasoning to trigger tool calls and fetch data in the background. This ensures an uninterrupted conversational flow, allowing the avatar to handle complex tasks like checking in a guest at a hotel without pausing the interaction.

  • Allows Customization to Match Brand Needs

    In addition to a library of diverse preset avatars, organizations can generate custom Live Avatars from high-quality reference images. This preserves brand styling and character identity, though custom avatar creation is currently available only through enterprise allowlisting.

Enhancements (2) ›
  • Enables More Natural, Multimodal Interactions

    By simultaneously processing visual and audio inputs, Live Avatar facilitates richer conversations that mirror human interaction. This capability allows enterprise agents to engage in dialogues incorporating listening, looking, speaking, and facial expressions for a more comprehensive experience.

  • Offers Seamless Multilingual Support Across 97 Languages

    The feature includes native multilingual speech-to-speech synchronization, dynamically adapting lip-sync and expressions. It can seamlessly transition across 97 languages without visual degradation or drift, providing natural conversational experiences at a global scale.

Notes (1) ›
  • Ensures Trust and Transparency with SynthID Watermarking

    All AI-generated audio and video output from Live Avatar is watermarked with SynthID, an imperceptible watermark woven directly into the content. This helps ensure AI-generated content remains detectable, minimizing misinformation and misattribution.

Read the original announcement →

https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-with-live-avatar/

Related releases