Google gave the AI avatar a face and didn't give it a price — that's the part to plan for
TL;DR: Gemini 3.8 Live with Live Avatar went generally available in Gemini Enterprise on 24 September 2026 — real-time lip-synced video over Google’s speech-to-speech stack, 97 languages, live camera and screen-share understanding, async tool calling, US and EU endpoints, SynthID watermarks in both the audio and the video. The conversation bills on the existing Live API meter: $3 per million audio input tokens, $12 per million audio output tokens, which at Gemini’s documented 25 audio-tokens-per-second is about $0.011 per conversational minute. The video bills at nothing — Google published no rate for it. HeyGen’s LiveAvatar charges $0.10/min in LITE mode for exactly that layer and $0.20/min in FULL mode; Tavus’s Growth tier works out to $0.32/min. The catches: Live Avatar has no standalone model ID and cannot be called outside Gemini Enterprise; custom avatars are allowlist-only; text and thinking tokens sit on a separate, accumulating meter; and an unpriced component on a GA enterprise product is a forward liability, not a discount — Google has already put dated January increases on two other products in this family.
What shipped
On 24 September 2026 Google moved Gemini 3.8 Live with Live Avatar to general availability in Gemini Enterprise, with endpoints in the United States and the European Union.
The headline feature is conversational video: the model generates a talking avatar in near real time, lip-synced to its own speech output. Organisations pick from a curated library of preset avatars, or — behind an enterprise allowlist and identity verification — generate a custom avatar from a high-quality reference image.
The rest of the 3.8 Live generation arrived with it. Language coverage is 97 languages with automatic detection, so a caller can switch mid-session without a configuration change. Interruption handling improved: the model recovers from a cut-off without losing conversational context, which was the specific failure that made earlier speech-to-speech agents unusable on real phone lines. Live visual understanding processes a camera feed or a screen share concurrently with the audio stream. Asynchronous tool calling lets the model fire a backend API and keep talking while it waits, rather than going silent for the duration of a lookup. Google is targeting web, mobile and interactive kiosks, with support in its Agent Development Kit for bidirectional streaming.
Every generated audio and video stream carries an imperceptible SynthID watermark. That is a meaningful extension — Google has marked generated audio since Lyria 3.5 reached the Gemini API, but marking the video track of a live-rendered human face is the first time the technique has been applied to a real-time synthetic likeness at this scale.
A separate Gemini 3.8 Live Extended Thinking variant, for agents that need multi-step reasoning mid-conversation, remains in private preview.
The number that isn’t there
Talking-head avatars are not new. Synthesia has sold them since 2021 and HeyGen has been undercutting it for nearly as long. Real-time conversational avatars — a face that listens and answers rather than reads a script — are newer but not novel; Tavus, D-ID and HeyGen’s LiveAvatar have all shipped one.
So the capability is not the story. The rate card is, and specifically the line that is missing from it.
Google prices Gemini 3.8 Live by modality:
| Modality | Rate |
|---|---|
| Audio input | $3.00 per million tokens |
| Audio output | $12.00 per million tokens |
| Text input | $0.75 per million tokens |
| Text output (thinking counted as output) | $4.50 per million tokens |
| Generated avatar video | no published rate |
Gemini’s documented audio tokenisation is 25 tokens per second. A minute of audio is therefore 1,500 tokens, in either direction:
- Listening: 1,500 × $3 ÷ 1,000,000 = $0.0045 per minute
- Speaking: 1,500 × $12 ÷ 1,000,000 = $0.018 per minute
In a turn-taking conversation where each party holds the floor roughly half the time, the audio cost of a minute is about $0.011. Even the implausible ceiling — billed a full minute of listening and a full minute of speaking — is $0.0225.
Now put the avatar market next to that. HeyGen’s LiveAvatar splits into two modes, and the split is what makes the comparison legible:
| Product | Mode | Effective rate |
|---|---|---|
| HeyGen LiveAvatar | FULL — HeyGen supplies LLM, ASR and TTS | |
| HeyGen LiveAvatar | LITE — you supply the AI stack | |
| Tavus | Growth tier, $397 for 1,250 minutes | ~$0.32/min |
| Gemini 3.8 Live + Live Avatar | audio metered, video unpriced | ~$0.011/min audio, $0 video |
LITE mode is the honest comparison. In LITE mode you are paying HeyGen $0.10 a minute for one thing only: rendering a face over an audio stream you generated and paid for somewhere else. That is precisely the service Google just attached to its own voice meter with no price on it.
This is not the pattern we saw a day earlier when Alibaba dropped speech-to-text through the price floor. Alibaba moved a published number by an order of magnitude and changed the unit under it. Google did something structurally different and, for a buyer, harder to respond to: it removed the number.
Why an unpriced feature is a liability, not a gift
There are two readings of a blank line on a rate card.
The generous one: video rendering is cheap enough at Google’s scale to absorb as a platform differentiator, the way intra-region storage egress is absorbed. Google has the TPU fleet and the incentive — Gemini Enterprise seats are the product, and the avatar sells seats.
The one a procurement team should actually use: you cannot sign against a rate that does not exist. No published number means no unit economic, no forecast, no contractual protection, and nothing preventing a rate appearing at the next terms update. That is not a paranoid reading of Google specifically. It is the pattern this product family has already established twice this month, in public, with dates attached: Gemini 3.8 Flash arrived carrying a shared January price cliff, and Gemini 3.8 Flash TTS shipped with a January doubling disclosed at launch.
Those two at least told you the number and the date. Live Avatar’s video gives you neither. The disciplined move is to budget the avatar at a plausible future rate — HeyGen’s $0.10-a-minute LITE tier is the obvious reference point for what the market thinks this layer is worth — and treat any saving against that as upside rather than baseline.
The three things that stop this being a swap
There is no API model ID. Live Avatar is a Gemini Enterprise feature, reachable through the Gemini Enterprise console and the Live API inside that platform. It is not a model a developer can point an arbitrary client at. If your product embeds a streaming avatar today through HeyGen’s or Tavus’s endpoint, there is currently no equivalent endpoint to migrate to. That single constraint defers most of the competitive pressure by at least a quarter.
Custom avatars are allowlisted. The preset library is available to all Gemini Enterprise customers. Generating an avatar from a reference image — your chief executive, your named support persona, your brand character — requires enterprise verification and a sales conversation. This matters because custom likeness is the high-value use case, and it is the one HeyGen and Synthesia have built self-serve for years. Google has chosen the slow, gated version of the feature that actually sells.
The audio meter is not flat. Unlike GPT-Live-1’s flat $0.05 per wall-clock minute, Gemini’s Live API bills accumulated context. Text and thinking tokens ride a separate meter at $0.75 and $4.50 per million, and a long session re-bills its growing context. The ~$0.011 audio figure is a floor for a short exchange, not a quote for a twenty-minute support call. That distinction is the same one that decides whether flat or token billing is cheaper for your call mix — dense conversations favour token meters, silence-heavy ones favour flat.
The compliance layer nobody markets
Article 50 of the EU AI Act has been in force since 2 August 2026, and a photorealistic conversational avatar touches three of its paragraphs at once.
Article 50(1) requires systems interacting directly with natural persons to make clear the person is interacting with an AI. An avatar that looks and sounds like a support agent is the textbook case.
Article 50(2) puts machine-readable marking on the provider. Google’s SynthID in both tracks is built for this, and it is genuinely useful: it is one obligation you are not implementing yourself.
Article 50(4) lands on the deployer — you. A deepfake, in the Act’s definition, is AI-generated or manipulated image, audio or video content resembling existing persons and appearing authentic. Deployers must disclose that such content is artificially generated, clearly and distinguishably, to the person exposed, at first exposure at the latest. An imperceptible watermark does not discharge this. You need a visible or audible label a viewer can perceive without any technical tool.
The preset-library avatar, which resembles no identifiable real person, sits further from the deepfake definition than a custom avatar built from your executive’s photograph. The allowlisted custom route is the one that pulls 50(4) squarely into scope — and it is, of course, the one most organisations actually want.
Which raises the release’s sharpest unanswered question. Google blocked voice replication in the EU entirely for Gemini 3.8 Flash TTS, one day before shipping face avatars with an EU endpoint. Cloning a likeness is not self-evidently a lighter obligation than cloning a voice. The most plausible reconciliation is the access model: voice replication was self-serve on a developer API, while custom avatars require named-account allowlisting and identity verification, and a contracted enterprise deployer is defensible in a way an anonymous API key is not. That is our inference from how the feature is gated, not a published rationale. If you plan to run a custom likeness on the EU endpoint, get Google’s written position on Article 50 allocation during allowlisting.
What to do
Already on Gemini Enterprise: run the preset avatar against a low-stakes internal workflow this week. Measure the bill across a realistic session length, not a two-minute demo — the text and thinking meter is where a long call’s cost actually lives. And test whether visual presence moves completion rates at all, because most evidence that it does is vendor-supplied.
Mid-procurement with an avatar vendor: do not cancel, but stop signing multi-year terms on the streaming line specifically. The pixels-only tier now has a credible free substitute in the pipeline. Pre-rendered video, custom avatar libraries and brand tooling are not under the same pressure — see best AI video tools for where the category currently stands.
Building voice agents with no video plans: nothing here changes your stack. Your decision remains the meter shape, and that turns on how much silence your calls contain. The best AI audio tools roundup and our ElevenLabs review cover the synthesis side; for customer-service agents specifically, Sierra and the Sierra vs Decagon comparison remain the relevant reference points.
Marketing and content teams evaluating avatars for outbound video should read this as a pricing signal rather than a product one — the marketers guide and our Gemini review cover the broader platform picture.
The one thing not to do is model the avatar at zero. Google put a face on its voice model and left the meter off. Meters get switched on.
Frequently asked questions
What exactly shipped, and can I call it from an API today?
Gemini 3.8 Live with Live Avatar went generally available in Gemini Enterprise on 24 September 2026, with endpoints in the United States and the European Union. The feature adds near-real-time generated video with synchronised lip-sync to Google's native speech-to-speech models. Alongside the avatar, the 3.8 Live generation carries 97-language support with automatic language detection, improved interruption recovery that does not drop conversational context, live visual understanding over a camera feed or screen share concurrently with audio, and asynchronous tool calling so the model can run a backend API in the background without pausing the conversation. It targets web, mobile and interactive kiosks. The availability caveat matters more than any of that: Live Avatar has no standalone model ID. It is a Gemini Enterprise feature, reachable through the Gemini Enterprise console and the Live API within that platform, not a model you can point an arbitrary client at the way you can point one at HeyGen's streaming endpoint. A separate Gemini 3.8 Live Extended Thinking variant remains in private preview. If your plan was to swap an avatar vendor out of an existing product this quarter, this release does not let you do that yet.
What does a minute of conversation actually cost?
On the audio meter, roughly one US cent, and you can check the arithmetic yourself. Google prices both gemini-3.8-live and the extended-thinking variant by modality: $3 per million audio input tokens, $12 per million audio output tokens, $0.75 per million text input tokens, and $4.50 per million text output tokens with thinking counted as output. Gemini's documented audio tokenisation is 25 tokens per second, so a minute of audio is 1,500 tokens in either direction. That gives $0.0045 per minute of listening and $0.018 per minute of speaking. A turn-taking conversation where each side holds the floor about half the time therefore lands near $0.011 a minute, and an implausible worst case where you are billed a full minute of both simultaneously is $0.0225. The honest caveat is that this is not a flat per-minute price the way OpenAI's GPT-Live-1 is. Text and thinking tokens sit on their own meter, accumulated context is re-billed as the session grows, and tool calls add their own traffic. The audio figure is a floor for a short session, not a quote for a long one. Model the text side separately for anything past a few minutes.
What does the avatar video cost?
Nothing, today, and that is the finding rather than the good news. Google has not published a rate for the generated video stream, because it is not sold as a separate line item — the avatar is presented as a feature of Gemini 3.8 Live rather than a product with a meter. There are two ways to read an unpriced component on a generally available enterprise product. The charitable reading is that video rendering is cheap enough at Google's scale to absorb as a platform differentiator, the way object storage egress inside a region is absorbed. The reading a procurement team should actually use is that an unpriced line is an unforecastable one: you cannot sign against a rate that does not exist, you cannot model a unit economic that has no unit, and nothing in a GA announcement prevents a price appearing at the next billing-terms update. Google has introduced dated price changes into this exact product family twice in the past month — the Gemini 3.8 Flash generation and Gemini 3.8 Flash TTS both carry January 2026 increases that were disclosed at launch. Here there is no disclosed number to plan around at all. Budget the avatar at a plausible future rate, not at zero.
Does this kill HeyGen, Synthesia and Tavus?
It attacks a specific part of their pricing, not their business, and the part it attacks is narrower than the headline suggests. HeyGen's LiveAvatar splits into two modes. FULL mode, where HeyGen supplies the language model, speech recognition and text-to-speech as well as the face, runs about $0.20 a minute — roughly $12 an hour — on the Essential tier with overage around $0.22 to $0.24. LITE mode, where you bring your own AI stack and HeyGen only renders the talking head, runs about $0.10 a minute, or $6 an hour, plus whatever your own model costs. Tavus's Growth tier is $397 a month for 1,250 conversational minutes, about $0.32 a minute. LITE mode is the direct comparison, because rendering a face over an audio stream you already paid for is exactly what Google is now doing at no marked price. That is a real competitive problem for the pixels-only tier. It is much less of a problem for the rest: HeyGen and Synthesia sell self-serve custom avatars, brand asset management, a stock avatar library measured in the hundreds, asynchronous pre-rendered video for training and marketing, and — critically — an API a developer can integrate this afternoon. Google's custom avatars are allowlisted behind enterprise verification, its avatar library is curated rather than deep, and there is no general API. See our [HeyGen review](/tools/heygen/) and [Synthesia review](/tools/synthesia/) for where each currently sits.
What are the EU AI Act implications of running a photorealistic avatar?
Two obligations, and Google discharges one of them for you. Article 50's transparency rules have applied since 2 August 2026. Article 50(1) requires that systems interacting directly with natural persons make clear the person is dealing with an AI — an avatar that looks and sounds like a human support agent is the paradigm case. Article 50(2) puts a machine-readable marking duty on the provider of the generative system, and Google's SynthID watermark, which it says is woven into both the audio and the video output, is aimed squarely at that. Article 50(4) is the one that lands on you: a deployer of a deepfake, meaning AI-generated video resembling an existing person and appearing authentic, must disclose that the content is artificially generated, clearly and distinguishably, to the person exposed to it, at first exposure at the latest. An imperceptible watermark does not satisfy that. You need a visible or audible label the viewer can perceive without any technical tool. A preset library avatar that resembles no real individual sits further from the deepfake definition than a custom avatar built from a reference image of your actual chief executive, and the custom route is the one that pulls Article 50(4) fully into scope. Our [Article 50 explainer](/news/eu-ai-act-article-50-live-what-changes-for-ai-tool-users-2026-08-02/) covers the general shape.
Google blocked EU voice cloning last week. Why is the EU getting a face avatar?
This is the sharpest open question in the release, and Google has not addressed it publicly. On 23 September, Gemini 3.8 Flash TTS shipped with consent-gated voice replication that Google blocked in the European Union outright rather than operate under Article 50 — a conservative call we read at the time as Google deciding the compliance surface of cloning a real person's voice in Europe was not worth the feature. A day later, Live Avatar shipped with a European endpoint and the ability to generate a custom avatar from a high-quality reference image of a person. Cloning a face is not obviously a lighter obligation than cloning a voice. The most plausible reconciliation is the gating: voice replication in Flash TTS was a self-serve capability on a developer API, whereas custom avatars require enterprise allowlisting, named-account verification and identity safeguards Google says are designed to prevent misuse. A contracted, verified enterprise customer is a defensible deployer in a way an anonymous API key is not. That is a reasonable design, but it is our inference from the access model, not a published rationale. If you intend to run a custom likeness on the EU endpoint, get Google's position on Article 50 allocation in writing during allowlisting rather than assuming the EU endpoint's existence settles it.
So what should I actually do this week?
Depends which side of the buy you are on. If you already run Gemini Enterprise, this is a console experiment, not a project — turn the preset avatar on against a low-stakes internal workflow, measure what the text and thinking tokens do to the bill over a realistic session length rather than a two-minute demo, and find out whether the visual presence changes completion rates at all. Most of the evidence that talking heads improve outcomes over plain voice is vendor-supplied. If you are mid-procurement with HeyGen, Synthesia or Tavus, do not cancel, but stop signing multi-year terms on the streaming line specifically; the pixels-only tier is the part with a credible free substitute arriving, and a twelve-month commitment priced against today's market is the wrong instrument. If you are building a voice agent and were never going to add video, nothing here changes your stack — [GPT-Live-1's flat $0.05 per minute](/news/openai-gpt-live-1-api-flat-per-minute-voice-meter-change-2026-09-10/) and Gemini's token meter remain the real decision, and that one turns on how much silence your calls contain.
Sources
- Google Cloud Blog — Gemini 3.8 Live with Live Avatar is now generally available
- Unite.AI — Google Brings Live Avatar Visual Presence to Gemini 3.8 Live
- Google Cloud — Gemini Enterprise Agent Platform pricing (Live API audio and text rates)
- Google Cloud — Live API documentation, Gemini Enterprise Agent Platform
- OrcaRouter — Gemini 3.8 Live with Live Avatar: Google's Enterprise Face (per-minute audio equivalents; video unpriced in public)
- ApiPulse — Gemini 3.8 Live API: Pricing, Thinking & Migration (modality-split rate card)
- Realtime Avatar — HeyGen API pricing explained (2026): video credits, the LiveAvatar split, and the realtime math
- Tavus — Plans and Pricing (conversational video minutes per tier)
- EU Artificial Intelligence Act — Article 50: Transparency Obligations for Providers and Deployers of Certain AI Systems
- European Commission — Transparency obligations under Article 50 of the AI Act (FAQ)
- Google ADK — bidirectional streaming documentation
Related tool reviews
Questions or corrections? Email Pick Right. Want the full list? See all news.