Google's new voice models undercut everyone — but the headline feature is unavailable in the EU, and the price doubles on 1 January
TL;DR: Google shipped Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS on 23 September 2026 at $0.50 per million text tokens in and $9.00 / $6.00 per million audio tokens out. Audio meters at 25 tokens per second, so an hour of speech is 81 cents on Flash and 54 cents on Flash-Lite — roughly a third of ElevenLabs. Three things the headline does not carry. Every rate doubles on 1 January 2027, the same date the Flash text line resets, which pushes Flash TTS above Cartesia Sonic 3.6 on cost. Voice replication is unavailable in the EEA, UK, Switzerland, Illinois, Texas and India. And Google’s #1 claim is measured on Hume AI’s benchmark — whose founder is now a Google DeepMind director. On the independent Voice Arena it is rank 2, inside the error bars.
What shipped
Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS reached general availability on 23 September 2026, logged in the Gemini API release notes the same week and replacing Gemini 3.1 Flash TTS Preview. Both are available today in the Gemini API and Google AI Studio, with Gemini Enterprise listed as coming soon. Flash TTS also reaches everyone through Gemini Notebook; Flash-Lite TTS through Google Vids.
The capability list is genuinely broad. Both models cover 100-plus languages and dialects — Google names Mexican Spanish, Quebec French, Scots English, Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic and Hindi among them — and ship with over 2,000 production-ready voices. Beyond the library there are three paths to a voice: design one from a natural-language description, replicate one from a 30-second sample, or remix timbre, pitch, pace and accent, which Google lists as coming soon rather than shipping.
The direction controls are the part that will matter most to anyone producing long-form audio. Scripts take line-by-line stage directions in natural language, generation runs to hours of continuous audio, two-speaker scenes can be staged in a single pass, and the models handle vocal bursts and backchannelling through inline markup such as <laughs> and |mhm|.
The rate card, and the date attached to it
| Model | Text in | Audio out | Per hour of speech |
|---|---|---|---|
| Gemini 3.8 Flash TTS | $0.50 | $9.00 | ~$0.81 |
| Gemini 3.8 Flash-Lite TTS | $0.50 | $6.00 | ~$0.54 |
| Flash TTS (Batch/Flex) | $0.25 | $4.50 | ~$0.41 |
| Flash-Lite TTS (Batch/Flex) | $0.25 | $3.00 | ~$0.27 |
| Flash TTS (Priority) | $0.90 | $16.20 | ~$1.46 |
| Flash-Lite TTS (Priority) | $0.90 | $10.80 | ~$0.97 |
Per million tokens, from Google’s own pricing page. The per-hour column follows from the audio meter: 25 tokens per second, so 90,000 tokens per hour.
Against the model it replaces, this looks like a large generational cut. Gemini 3.1 Flash TTS Preview charged $1.00 and $20.00, which is $1.80 for that same hour. Flash TTS at 81 cents is 55% cheaper.
Except the pricing page carries a second line for every one of those rates: through December 31, 2026. On 1 January 2027, Flash TTS goes to $1.00 and $18.00, Flash-Lite to $1.00 and $12.00. An hour of Flash TTS audio becomes $1.62 — against the $1.80 its predecessor charged. The durable generational saving on the audio meter is 10%. The other 45 points are a promotion with a date on it.
That date should look familiar. It is the same 31 December 2026 expiry that the Gemini Flash text models carry, where Gemini 3.8 Flash launched at the identical introductory rate, with the identical expiry, that 3.7 Flash had three weeks earlier. The pattern holds on the speech line: the discount follows the calendar, not the model. Standardising on the newest TTS model does not buy a single extra day of runway, and neither does waiting for the next one.
Where the price stops winning
Normalised to characters — the unit the rest of this market sells in — the current rates are about $33 per million characters for Flash TTS and $22.10 for Flash-Lite. Cartesia Sonic 3.6 is $49. ElevenLabs Eleven v3 is $100, at $0.10 per thousand characters.
On those numbers Google is the obvious buy, and the quality gap points the same way: on the Artificial Analysis Provider Voice Arena, Gemini 3.8 Flash TTS leads Eleven v3 by 93 Elo points, a margin with no overlap in the confidence intervals.
Run the same table at January rates and it reorders. Flash TTS doubles to roughly $66 per million characters — above Cartesia’s $49. And Cartesia is not a weaker model: on that same Voice Arena, Sonic 3.6 holds rank 1 at 1,273 Elo against Flash TTS at 1,260, with ±17-point intervals on both. Google’s model is not beating it now; in January it would also cost a third more.
Flash-Lite is the version that survives the reset. At about $44.20 per million characters it stays under Cartesia, and at rank 6 with 1,235 Elo it is a real step down in quality but not a collapse. For high-volume, less voice-critical work — the content pipelines where an hour of narration is a commodity input — that is the line item to model.
The benchmark provenance problem
Google’s announcement leads with rankings: #1 on Hume AI’s Voice Design Benchmark at 71.4, and the top two places on Hume’s Overall Quality Index.
Those are Google-reported results about Google models, which is ordinary for a launch. The part that is not ordinary is who maintains the benchmark. Alan Cowen, Hume AI’s founder, is credited on this release as a Director of Research Science at Google DeepMind, having joined through a January licensing deal that moved Hume engineers to Google while leaving Hume free to serve other labs.
Nothing here suggests the benchmark was manipulated. But a vendor announcing a first-place finish on a yardstick built by a company whose founder now works for that vendor is not an independent result, and it should not be read as one. The independent reading is the blind-preference Voice Arena, and it says something more modest: statistical tie for first, not a clear win. This is the same discipline that applies to independently repriced intelligence indices — the leaderboard a vendor chooses to headline is itself a claim.
Consent, watermarking and the EU exclusion
Every generated clip carries a SynthID watermark, and voice replication output additionally carries C2PA credentials. That matters for the same reason it mattered when Google shipped Lyria 3.5 with SynthID on every output: Article 50(2) of the EU AI Act puts a machine-readable marking duty on the provider, and Google meeting it before the file reaches you removes one compliance question from your build. It does not remove the disclosure duties that sit on you as a deployer — the same split that applies to watermarked text output.
To replicate a voice, Google requires a verbal consent recording from the voice owner that matches the reference speaker. That is a meaningful gate, and it is more than most of this market asks for.
It is also, for a large share of readers, beside the point. Voice replication is unavailable in the EEA, the United Kingdom, Switzerland, India, Illinois and Texas. Google built the consent verification, added C2PA on top, and still did not ship the feature into those jurisdictions. That decision tells you how the compliance exposure is being priced: not as something a consent recording clears, but as something worth forgoing the market over.
For an EU buyer the practical consequence is simple. The feature the launch coverage is organised around is not on your menu. The voice library, voice design, stage directions and long-form generation all are — and those, not cloning, are what most production pipelines actually run on.
What to do with this
If you are generating narration at volume, Google is now the cheapest credible option in the AI audio tool field, and Flash-Lite is the version whose economics survive January. Cost the workload at the post-reset rates before you commit; the promotional numbers are not what you will be paying by the time a pipeline built this quarter is running at scale.
If you are building conversational voice rather than generating audio files, this is the wrong meter to compare. OpenAI’s GPT-Live-1 bills wall-clock seconds rather than audio tokens, which changes the cost of silence, hold queues and thinking pauses entirely. And if the workload is speech going the other direction, the transcription price floor is a separate calculation again.
If your entity or your users sit in the EEA or the UK, treat voice replication as unavailable and design without it. The rest of Gemini’s speech stack is fully available — and on the evidence of the independent leaderboard, the prebuilt and designed voices are where the quality claim actually holds up.
Frequently asked questions
How much does an hour of Gemini 3.8 TTS audio actually cost?
Audio output bills at 25 tokens per second, so an hour of generated speech is 90,000 audio tokens. At the current promotional rate of $9.00 per million audio tokens, Gemini 3.8 Flash TTS costs about 81 cents per hour; Flash-Lite TTS at $6.00 works out at about 54 cents. Text input is a rounding error at $0.50 per million tokens — a script long enough to fill an hour of speech is a few thousand tokens, well under a cent. Batch and Flex tiers halve the audio rate to $4.50 and $3.00, putting an hour at roughly 41 and 27 cents for work that tolerates delay; the Priority tier raises it to $16.20 and $10.80, or about $1.46 and 97 cents. All of these are introductory rates. From 1 January 2027 every one of them doubles, taking standard-tier Flash TTS to about $1.62 per hour and Flash-Lite to about $1.08.
Can EU buyers use the voice replication feature?
No. Google lists voice replication as unavailable in the European Economic Area, the United Kingdom, Switzerland, India, and the US states of Illinois and Texas. This is a hard geographic exclusion on the feature, not a consent workflow that EU customers can satisfy — Google built the consent machinery, which requires a verbal consent recording from the voice owner that matches the reference speaker, and still withheld the capability from those markets. Everything else in the release is available: the 2,000-plus prebuilt voices, voice design from a natural-language description, line-by-line stage directions, two-speaker staging and long-form generation all work normally. Only cloning a specific real person's voice from a 30-second sample is fenced off. For an EU team, this means the headline capability in the launch coverage is not part of what you are buying, and any build plan written from the announcement needs that line removed before it is costed.
Is Gemini 3.8 Flash TTS really the best voice model available?
Google's claim is that it ranks #1 on Hume AI's Voice Design Benchmark with a score of 71.4, and takes the top two places on Hume's Overall Quality Index. Two caveats belong next to that. First, those are results Google reported about its own model. Second, Alan Cowen — Hume AI's founder — is credited on the release as a Google DeepMind Director of Research Science, having joined through a January licensing deal. Google topping a benchmark maintained by a company whose founder it has since hired is not evidence of wrongdoing, but it is not independent either. The independent number is the Artificial Analysis Provider Voice Arena, a blind preference leaderboard. There, Gemini 3.8 Flash TTS sits at rank 2 with an Elo of 1,260 against Cartesia Sonic 3.6 at 1,273. Both carry ±17-point confidence intervals over roughly 1,750 to 2,000 samples, so the honest reading is a statistical tie for first, not a win. Flash-Lite TTS is rank 6 at 1,235.
How does the price compare to ElevenLabs and Cartesia?
Normalised to a per-million-characters basis, Gemini 3.8 Flash TTS is about $33 and Flash-Lite about $22.10, against $49 for Cartesia Sonic 3.6 and $100 for ElevenLabs Eleven v3 at $0.10 per thousand characters. On today's rates Google is roughly a third of ElevenLabs and a third cheaper than Cartesia, while outranking ElevenLabs by 93 Elo points on the Voice Arena. The comparison inverts on 1 January 2027. When Google's promotional rates lapse, Flash TTS goes to roughly $66 per million characters — more than Cartesia's $49, against a model that already edges it on the independent leaderboard. Flash-Lite lands at about $44.20 and stays under Cartesia. Note that the vendors meter differently: ElevenLabs and Cartesia charge per character, Google per audio token, and Cartesia's subscription structure means the marginal cost of additional characters is effectively zero until you hit the concurrency cap.
What should I do before committing a pipeline to these models?
Three things. Cost the workload at the January rates rather than today's, because the introductory pricing runs out on 31 December 2026 and nothing about adopting the newest model extends it — this is the same shared expiry that governs the Gemini Flash text models, and it attaches to the calendar rather than to the version you pick. Check whether your build actually depends on voice replication, and if your users or your entity sit in the EEA, the UK or Switzerland, design around its absence now rather than discovering it at integration. And confirm which tier you are billed on: the gap between Batch at $3.00 and Priority at $16.20 per million audio tokens on Flash-Lite is more than five to one, which is a far larger lever than the choice between Google and its competitors.
Sources
- Google — Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS
- Google AI for Developers — Gemini API pricing
- Google AI for Developers — Gemini API release notes
- OrcaRouter — Gemini 3.8 Flash TTS: Google's speech line splits in two
- OrcaRouter — Gemini 3.8 TTS vs Cartesia Sonic 3.6: Elo vs price
- OrcaRouter — Gemini 3.8 TTS vs ElevenLabs: what 3x the price buys
- RuntimeWire — Google launches Gemini 3.8 voice models with Hume founder Alan Cowen credited
Related tool reviews
Questions or corrections? Email Pick Right. Want the full list? See all news.