Lilith Lilith.
⌕
Editorial illustration: Gemini 3.8 Live adds a face that speaks 97 languages
Lilith illustration · editorial remix

Google has combined the live dialogue of Gemini 3.8 Live with continuously generated avatar video. An enterprise agent can listen, process visual input, speak and change expression throughout a conversation without waiting for every tool call to finish.

Live Avatar joins speech, video and tools in one stream

Gemini 3.8 Live with Live Avatar generates audio and video in near real time. Google describes precise lip sync, fluid turn taking and natural expressions. Asynchronous function calling lets the agent fetch data or run a tool in the background while the conversation continues.

The system can switch among 97 languages and adapt its mouth movements and expressions as it goes. Google provides a library of preset characters. A high quality reference image can also produce a custom avatar that preserves a brand or character identity, although this option is limited to allowlisted enterprise customers. Google says all generated audio and video carries a SynthID watermark.

The enterprise agent gains a social interface, not merely another output format

For customer support, onboarding and interactive walkthroughs, a face becomes a functional layer. Users can see a reaction while the backend is working and may better understand when the system is listening or responding. Asynchronous tools may matter more than the graphics because they keep the exchange alive during background work.

Product teams also gain a much larger surface to design and test. Answer accuracy is now joined by voice, expression, lip sync, latency and tool state. A failure is no longer just a bad sentence. It can be a confident smile while a payment has been declined.

A fluid face can conceal an inaccurate or expensive backend

The announcement provides no public measurements for end to end latency, lip sync error, the price of the video layer or performance under sustained load. Availability through Gemini Enterprise does not make it an open consumer product, and custom avatar identity remains restricted by allowlisting.

SynthID helps establish provenance, but it does not resolve consent for a reference image, access control or answer correctness. Enterprise deployments will still need logs, clear AI disclosure and a reliable path to a human.

Latency, session cost and failure behavior will determine its value

The first meaningful signals will come from long customer conversations: whether audio and video drift apart, whether a background tool stalls the session and what a call lasting several minutes costs. How the avatar communicates uncertainty or a technical failure will matter just as much.

Google has demonstrated a persuasive presentation layer. Production value will require measurements across the whole path from microphone through function calling to video, not simply a convincing face.

Lilith's verdict

Google has placed a face behind the counter that speaks 97 languages and keeps talking while tools run. An enterprise will value it when the smile stops during a failure and the right human is called.

I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.

Original source ↗ ↗