Google releases Gemini 3.8 Live

25-09-2026

Google made Gemini 3.8 Live with Live Avatar generally available on 24 September 2026. It gives voice agents an animated, lip-synced face across 97 languages.

Written by:

Senne Doets

Online Marketeer at DataNorth | Next-Gen AI & Tech Apprentice

google releases gemini 3.8 live
Sign up for our Newsletter

Published: 25 September 2026

Google made Gemini 3.8 Live with Live Avatar generally available on 24 September 2026. It gives a voice agent a face: an animated presenter that speaks and lip-syncs in 97 languages, billed at $1.00 per million video output tokens on top of the audio you already pay for. It runs in Gemini Enterprise only, on US and EU endpoints.

What does Gemini 3.8 Live with Live Avatar do?

It turns a spoken conversation into a video call with a synthetic presenter. The model takes speech in and returns speech and video out, so the face you see is generated in the same pass as the voice. There is no separate lip-sync step bolted on afterwards.

The avatar speaks 97 languages and switches between them on its own, without you telling it which one the caller used. It can also see. The agent takes camera feeds and screen shares as input, so a support session can watch what the customer is pointing at.

You pick the face in one of two ways. Google ships a library of ready-made avatars that any Gemini Enterprise customer can use today. Building an avatar from your own reference photo needs an allowlist and a verification step, so it is not generally available.

What comes with it:

  • Native speech to speech, so there is no speech-to-text step in the middle
  • Tool calling that runs in the background while the avatar keeps talking
  • Interruption recovery, so a caller can cut the avatar off mid-sentence
  • SynthID watermarks on every second of audio and video it produces
  • US and EU endpoints, with provisioned throughput for steady traffic
  • Gemini 3.8 Live Extended Thinking is still in private preview, not generally available

Gemini 3.8 Live with Live Avatar pricing

Take this from the table: the face is cheap per token and expensive per minute, because video burns roughly 250 times more tokens per second than audio.

What is measuredWith Live AvatarAudio only
Text input, per million tokens$0.75$0.75
Audio input, per million tokens$3.00$3.00
Audio output, per million tokens$12.00$12.00
Video output, per million tokens$1.00not billed
Tokens per second of output6,192 video plus 25 audio25 audio
One minute of the agent speakingabout $0.39about $0.02

The per-token prices come from Google’s Gemini Enterprise pricing page, not from the launch post, and they are introductory rates that run to 31 December 2026. From 1 January 2027 the standard rate of $1.50 in and $7.50 out per million tokens applies. The two per-minute figures in the last row are ours, worked out from Google’s own tokens-per-second rates. Google bills video only while the avatar is speaking, so listening time is free.

What Google is not saying about Live Avatar

There is no latency figure. For a live video agent, time to first frame is the number that decides whether the thing feels like a conversation or a slideshow, and Google calls it near real-time without publishing a millisecond anywhere.

There is also no benchmark and no evaluation of any kind. The launch post names customers, including Salesforce Agentforce and Cox Automotive, but it does not say how often the avatar recovers correctly from an interruption, how well the lip-sync holds across the 97 languages, or which of those languages were tested at all. Ninety-seven is a count of supported languages, not a quality claim.

And custom avatars, the feature most enterprises actually want, sit behind an allowlist. What is generally available today is a library of Google’s faces, not yours.

What this means

Worth testing now if you already run Gemini Enterprise and have a customer-facing flow where a face changes the outcome: onboarding walkthroughs, guided troubleshooting, or product demos in markets where you cannot staff native speakers. At around $0.39 a minute of speaking time, a ten-minute guided session costs about four dollars. That is cheaper than the person it replaces and far more expensive than the chatbot it replaces, and the band between those two is narrower than it first looks. Measure time to first frame yourself on your own traffic, because Google has published nothing you can plan against.

Safe to ignore for now if you sit outside Gemini Enterprise, or if what you wanted was your own brand’s face on screen. The allowlist decides that, not your budget. The wider point is that Google has moved avatar generation out of a standalone product and into the same call as the voice, which is direct pressure on Synthesia, HeyGen and D-ID. Their pitch was that you bolt an avatar onto a model. Google’s answer is that you no longer need to, and whether the seam is really gone is exactly what the missing latency figure would tell you.

For more information, visit the official announcement of Gemini 3.8 Live with Live Avatar on the Google blog.

Add DataNorth AI to your Google favorites