Primary sourceAIArticle··4 min read

Gemini 3.8 Live Gets a Face: Google Launches Live Avatar

After voice, Google is giving its conversational agents a video avatar synced in near real-time, designed for enterprise customer service.

Gemini 3.8 Live Gets a Face: Google Launches Live Avatar
Source : Shuo-yiin Chang et CJ Zheng (Google) · Google · 24 September 2026View original ↗

In brief

Google adds Live Avatar to Gemini 3.8 Live: a video avatar generated in near real-time, natively coupled with voice dialogue, with lip-sync in 97 languages and asynchronous tool calls. Available now in Gemini Enterprise, it targets customer service and interactive journeys. One more step toward AI agents that resemble human interlocutors, with SynthID as a safeguard.

🍺 Bar-stool version

Google looked at its voice assistants and decided the problem was that you couldn't see their face. So now your hotel's agent smiles at you, moves its lips in 97 languages, and checks your reservation while making small talk, like a receptionist who never sleeps or needs a coffee break. The whole thing is tattooed with an invisible watermark, just so we can prove afterward that this charming gentleman doesn't exist. The front desk is getting a face, it just won't be anyone's.

Key takeaways

  1. 1

    Live Avatar natively couples Gemini 3.8's live dialogue model with low-latency streaming video generation.

  2. 2

    The avatar listens, sees, and speaks: it processes audio and visual input simultaneously and responds with voice, lip-sync, and facial expressions.

  3. 3

    Asynchronous tool calls let the agent query data in the background without interrupting the conversation, demonstrated with a hotel check-in.

  4. 4

    Lip-sync adapts to 97 languages, with mid-conversation language switches and no reported visual drift.

  5. 5

    Businesses have access to prebuilt avatars and can create a custom avatar from a single reference image, but only via allowlisting.

  6. 6

    All audio and video outputs are marked with Google's SynthID watermark.

  7. 7

    The feature is available starting September 24, 2026 in Gemini Enterprise, one week after the launch of Gemini 3.8 Live.

A Face for Gemini Live

One week after launching Gemini 3.8 Live, Google is adding a video layer to its real-time dialogue model. Live Avatar generates an animated character that talks, with lip-sync, natural expressions, and smooth turn-taking.

The announcement, signed by researcher Shuo-yiin Chang and engineer CJ Zheng on behalf of the Gemini Audio team, emphasizes the 'native' coupling between dialogue and streaming video. In other words, the video isn't a wrapper slapped on top of a synthetic voice after the fact.

The target is clearly enterprise: customer service, interactive guided tours, and more broadly any virtual offering a brand might want to make more embodied.

See, Hear, Respond

Google starts from a simple observation: a conversation is multimodal. We listen, we look, we speak, and we express ourselves with our face.

Live Avatar processes what it sees and hears in parallel, then responds with expressive audio and video. Google calls it 'near real-time,' without giving a latency figure in the post.

Tools Called Without Cutting the Conversation

The most concrete point for integrators is asynchronous tool execution. The avatar can trigger an API call or fetch data in the background while continuing the exchange.

The demo Google chose is a hotel check-in: the agent verifies the reservation while still talking to the guest. This is what separates a decorative avatar from an agent capable of handling a real task without awkward silence.

97 Languages and Custom Avatars

Lip-sync and expressions adapt dynamically to 97 languages, with switches mid-conversation. Google claims these transitions happen without loss of video fidelity or visual drift.

On the identity side, a library of prebuilt avatars is provided. Businesses can also generate an animated avatar from a high-quality reference image, preserving likeness, brand identity, or a character's persona.

This customization remains restricted to allowlisted clients, a filter that is likely not unrelated to impersonation risks.

SynthID as a Safeguard

All audio and video outputs are marked with SynthID, Google's imperceptible watermark, embedded directly in the generated content. The stated goal is to limit misinformation and misattribution errors.

Google points to a model card for details on its safety approach, and to API documentation to get started.

“Live Avatar creates an experience that listens, sees, and speaks with a dynamic visual persona.”
“Conversation is inherently multimodal: we listen, look, speak, and use facial expressions to communicate.”
“Custom avatar creation is currently available only through enterprise allowlisting.”

Why it matters

Live Avatar marks the shift from voice agent to embodied agent, and Google is selling it first to businesses, where call center money lives. The most strategic combination isn't the face itself but the asynchronous tool calling: that's what makes an avatar useful rather than merely impressive. Several gray areas remain. The post gives no latency figures, no quality metrics, no pricing indication, and 'near real-time' can cover very different realities. Above all, creating avatars from a simple photo raises the deepfake question head-on: allowlisting and SynthID are answers, but a watermark only helps if someone checks it, and the customer facing the screen won't. The real question becomes transparency toward the end user, which Google largely leaves to its enterprise clients.

#google#gemini#avatar#agents#multimodal#enterprise
Original source
Introducing Gemini 3.8 Live with Live Avatar
Shuo-yiin Chang et CJ Zheng (Google)
Open the article ↗

For you

Put it to work on your sources.

Free: this week's articles and three sources of your own. Pro: the whole archive and your sources, from €8/month.

For your team

The same machine, on your topics.

A space in your colours, your watch angles, your curators. Pilot open to three companies.

Read next