AI-generated content. This article was researched and written by an automated AI editorial system and published without prior human review. Every factual claim is checked against cited primary sources before publication, but no journalist read this page before you did — treat it accordingly, and report anything that looks wrong. How this works ›

Some links on this page are affiliate links. We may earn a commission at no extra cost to you.
Updated: Aug 28, 2026
·
googlegeminispeech-to-texttranscriptionpricingdeepgramassemblyaiwhisperprocurementaudio

Google's new transcription model didn't reset the speech-to-text price floor — it finally brought Google down to it

TL;DR: Google put Gemini 3.5 Transcribe into public preview on 26 August 20262.6% average WER non-streaming and 4.0% streaming as measured by Artificial Analysis, 85+ languages with mid-session code-switching, and a 70% improvement in time-to-final-transcription over Chirp 3. It is being covered as Google undercutting the transcription market. It isn’t. At roughly $0.005/minute blended for batch — about $0.30 per audio-hour with diarization included — it lands above AssemblyAI ($0.23/hr loaded) and Deepgram Nova-3 ($0.28/hr loaded). The specialists set this floor years ago; Google just walked down to it. The big number is internal: Google Cloud Speech-to-Text costs $0.48/hr plus $0.36/hr for diarization — so Google just repriced its own transcription stack by about 64%. For you: on Google Cloud STT, migrate — the case is overwhelming. On a specialist, price is not a reason to move, and the cross-vendor accuracy numbers you are being shown are not comparable. And note that Gemini bills per token, not per audio-minute, so ~40% of your bill moves with how densely people talk.

What actually shipped

Two model IDs, both in public preview through the Gemini API and Google AI Studio:

Google’s stated capabilities are specific enough to check. Automatic detection across more than 85 languages, handling regional accents and dialects, with code-mixing supported mid-session. Smart transcription handles self-corrections, strips filler words, and auto-formats output. Custom vocabulary biases recognition toward domain jargon and unusual spellings, capped at 1,000 terms. Speaker diarization and word-level timestamps for pre-recorded audio. And function calling, which lets the transcription model hand off to other Gemini models mid-task.

It is already shipping inside Google products: Gboard Rambler on Android, the Gemini app for macOS, and Google Antigravity’s prompt box, with Chrome named as next. Enterprises get it in public preview through the Gemini Enterprise Agent Platform.

That is a substantial release. The question is what it does to the market, and the answer is not the one in the headlines.

The comparison that matters

Nearly every writeup of this launch quotes Google’s per-minute figure without loading in the add-ons that competitors charge separately — which is precisely the mistake that makes speech-to-text procurement go wrong. The relevant number is what you pay per audio-hour for pre-recorded transcription with speaker labels, because almost nobody transcribing meetings or calls needs it without them.

VendorBase (batch)DiarizationLoaded
AssemblyAI (async)$0.21/hr+$0.02/hr$0.23/hr
Deepgram Nova-3~$0.26/hr+$0.02/hr$0.28/hr
Gemini 3.5 Transcribe~$0.30/hrincluded$0.30/hr
OpenAI Whisper API~$0.36/hrnot available$0.36/hr
Google Cloud STT (Chirp)~$0.48/hr+$0.36/hr$0.84/hr
AWS Transcribe~$1.44/hr+$0.14/hr$1.58/hr

Google arrives third, not first. It is roughly 30% more expensive than AssemblyAI and 7% more than Deepgram for the same job.

The streaming picture is slightly worse. Gemini’s live endpoint blends to about $0.009/minute, or $0.54/hour, against roughly $0.45/hour for AssemblyAI streaming and $0.46/hour for Deepgram. On real-time transcription, Google is the expensive option among the three — and, as covered below, the one that cannot attribute speakers while streaming.

None of this makes the release unimpressive. It makes the framing wrong. The speech-to-text price floor was not moved this week. It was set by the specialists well before this launch, and Google has now met it rather than broken it.

The 64% that did happen

Look one row further down that table and the actual repricing is obvious.

Google Cloud Speech-to-Text — the Chirp-based product Google has been selling to enterprises — costs $0.48/hour for batch, and diarization is a $0.36/hour add-on, which is eighteen times what Deepgram and AssemblyAI charge for the same feature. Loaded, that is $0.84/hour.

Gemini 3.5 Transcribe does the same job, more accurately, at $0.30/hour with diarization bundled. That is a roughly 64% reduction in Google’s own transcription pricing, and it comes with the 70% latency improvement over Chirp 3 that Google leads with.

So the story is not Google attacking Deepgram. It is Google retiring an uncompetitive product it had been charging enterprise rates for, and folding transcription into the Gemini platform where the pricing is set by what the market already charges. That is a meaningful event — it is just an internal one, and it tells you something about where the pressure came from. Google did not price this against the specialists to win share from them; it priced it at the floor because that floor is now simply what speech-to-text costs, and continuing to charge $0.84/hour against it had become untenable.

Where the accuracy claim holds and where it breaks

The 2.6% figure is better evidence than most launch benchmarks because it comes from Artificial Analysis rather than Google’s internal evaluation. Independent measurement is worth something, and the pattern of vendors marking their own homework makes third-party numbers genuinely more credible.

What that number cannot do is rank Google against its competitors, because word error rate is close to meaningless across different test sets. The published figures currently circulating — Deepgram at 5.26%, AssemblyAI Universal-2 at ~14.5%, ElevenLabs Scribe v2 at ~3.3%, Google at 2.6% — are measured on four different corpora with different audio quality, accent distributions, domain vocabularies, and disfluency-scoring conventions. Deepgram’s number is deliberately drawn from hard real-world medical, finance and call-centre audio. ElevenLabs’ is English-only internal evaluation. Ranking them produces a leaderboard that means nothing.

The clearest demonstration is inside Google’s own release: the same model scores 2.6% on the Artificial Analysis set and 5.04% on FLEURS multilingual. One model, one week, and the accuracy claim nearly doubles depending on what you play it.

Run your own audio. A few hours of representative recordings, scored against a corrected transcript, will tell you more than every published WER in this category combined.

The limits that decide whether you can use it

These are in the documentation rather than the announcement, and they are the details most likely to disqualify the model for a given application:

Add to that the ordinary preview caveat: this is public preview, not GA, and preview pricing on Gemini models has a history of being introductory. Assume the rate can move before it stabilises.

The billing model nobody is mentioning

Every specialist in this category bills a flat rate per audio-minute. Google bills per token, and publishes a per-minute conversion.

For audio input that distinction does not matter — audio tokenises at a constant rate per second, so a minute of silence costs the same as a minute of dense speech. But the text output side is billed on the transcript you actually generate, at $12.00 per million tokens for batch. Four people talking quickly for an hour produce substantially more output than a sparse one-on-one with long pauses, and you pay the difference.

That variable component is roughly 40% of the blended batch estimate and 44% of the live estimate. Which means Deepgram’s $0.26/hour is a quote and Google’s $0.30/hour is an estimate. On a fixed small volume this is noise. On a large deployment, or anywhere you have to forecast spend rather than simply pay it, it is a real structural difference — and it runs in the wrong direction, because the workloads with the densest speech are usually the high-value ones you most want to transcribe.

What this changes for what you buy

If you buy meeting software, nothing. Otter.ai, Fireflies.ai and Granola do not compete on raw transcription accuracy — they compete on calendar integration, bot join behaviour, searchable cross-meeting history, CRM sync and permissions. Speech-to-text is one upstream component, and when it improves your vendor adopts it and you notice better transcripts. Our audio tools guide is unaffected by this release, and so is the choice between those three.

If you call an STT API directly, your position depends entirely on where you are now:

If you are building voice products more broadly, this sits alongside the realtime voice model consolidation as the same underlying trend: the audio layer of the stack is being absorbed into the general-purpose frontier platforms, and the standalone-component vendors are being pushed toward the parts the platforms have not bothered with yet. Real-time diarization is currently one of those parts. It will not stay one indefinitely.

The read

The interesting thing about Gemini 3.5 Transcribe is not that it is cheap. It is that Google could not price it above the specialists even with a frontier lab’s accuracy and its own distribution into Chrome, Gboard and Android.

Transcription has become genuinely commoditised. The floor is set by companies far smaller than Google, it did not move this week, and the largest AI company on earth arriving with better accuracy still had to meet it rather than beat it — while cutting its own legacy product by nearly two-thirds to do so.

For buyers that is good news of a specific kind: the price of speech-to-text is now stable, competitive, and low enough that it should not be the deciding factor in what you build. Pick on latency, on session limits, on whether you need speakers labelled in real time, and on what your own audio actually scores. Those still differ meaningfully between vendors.

The price no longer does.

Frequently asked questions

Is Gemini 3.5 Transcribe cheaper than Deepgram or AssemblyAI?

No, and this is the most common misreading of the launch. Compare loaded prices — base rate plus the add-ons you actually need — for pre-recorded audio with speaker labels. AssemblyAI is $0.21/hour async plus $0.02/hour for diarization, so $0.23/hour. Deepgram Nova-3 is roughly $0.26/hour batch plus $0.02/hour for diarization, so $0.28/hour. Gemini 3.5 Transcribe is about $0.30/hour with diarization included in the base rate. Google is the most expensive of the three, though the gap is small enough that it will not decide a mid-sized deployment on its own. Where Gemini does win on price is against OpenAI's Whisper API at roughly $0.36/hour, which additionally has no speaker diarization at all, and against the legacy enterprise options — Google's own Cloud Speech-to-Text at $0.84/hour loaded and AWS Transcribe at roughly $1.58/hour loaded. The honest summary: Google has joined the cheap tier rather than created a new one. If you are already on a specialist, price is not your reason to migrate.

Can I trust the 2.6% word error rate claim?

Trust it as a real measurement and distrust it as a comparison. The 2.6% non-streaming and 4.0% streaming figures come from Artificial Analysis, a genuine independent evaluator, which is meaningfully better than a vendor-only benchmark. The problem is that competing WER numbers come from different test sets, and word error rate is extraordinarily sensitive to test-set composition — audio quality, accent distribution, domain vocabulary, and how disfluencies are scored all move the number by several points. Deepgram's published 5.26% batch figure is measured on a deliberately hard real-world set spanning medical, finance and call-centre audio. AssemblyAI's ~14.5% comes from a challenging mixed set. ElevenLabs Scribe v2 reports ~3.3% on internal English-only evaluation. Lining those four numbers up in a table and declaring a winner is meaningless, because they are not measuring the same thing. Google's own multilingual figure demonstrates the effect within a single vendor: on the FLEURS benchmark the same model scores 5.04% non-streaming rather than 2.6%. The only number that should decide a purchase is the one you produce yourself by running a few hours of your own representative audio through the candidates and scoring the output against a corrected transcript.

Why does the token-based billing matter if Google publishes a per-minute price?

Because only half of the per-minute price is actually fixed per minute. Google quotes batch transcription as $2.00 per million audio input tokens, which it converts to roughly $0.003 per minute, plus $12.00 per million text output tokens, which it converts to roughly $0.002 per minute. The audio input side is genuinely fixed — audio is tokenised at a constant rate per second regardless of what is on the recording, so silence costs the same as speech. The output side is not fixed. You are billed for the transcript you generate, so a dense technical panel discussion with four people talking quickly produces far more output tokens per audio-minute than a sparse one-on-one call with long pauses. That variable component is about 40% of the blended batch estimate and about 44% of the live estimate. Every specialist vendor in this category bills a flat rate per audio-minute regardless of speech density, which means their quote is a quote and Google's is an estimate. For a fixed-volume workload this is a rounding error. For a large deployment, or for anyone who has to forecast a budget rather than just pay a bill, it is a genuine structural difference — run your own densest audio through it before extrapolating from the published per-minute figure.

What are the session limits, and will they break my use case?

These are the details most likely to disqualify the model for a specific application, and they are easy to miss. For pre-recorded audio the limit is one hour per request, but that drops to 30 minutes when you enable speaker diarization or word-level timestamps — which is to say, the limit is 30 minutes for essentially every meeting-transcription use case, so anything longer needs chunking and rejoining. Custom vocabulary is capped at 1,000 terms. The live streaming endpoint is limited to 10 minutes per session, so continuous transcription of a long call requires session rotation and reconnection handling that you have to build and test. Most importantly, speaker diarization and word-level timestamps are not available in live mode at all — they are file-processing features only. If you are building anything that needs real-time speaker attribution, such as live captioning that labels who is talking, Gemini 3.5 Transcribe does not currently do it and a specialist streaming vendor still does. There is also a documentation discrepancy worth confirming against your own testing: the API documentation states diarization supports up to eight speakers while Google's announcement blog describes accurate attribution for up to three.

Should we switch our meeting-notes tool because of this?

No. This is an API-layer release, not a product release, and the distinction matters for how you should react. Tools like Otter.ai, Fireflies.ai and Granola are not primarily selling you transcription accuracy — they are selling calendar integration, meeting-bot join behaviour, searchable history across months of calls, CRM sync, speaker-attributed summaries, and a permissions model. Raw speech-to-text is one upstream component of that, and if it gets cheaper or more accurate the vendors will adopt it quietly and you will notice nothing except gradually better transcripts. Nothing about this launch changes which meeting tool fits your team. The people who should act on this release are those calling a speech-to-text API directly in their own product — and among them, specifically those currently on Google Cloud Speech-to-Text, AWS Transcribe, or self-hosted Whisper without diarization, where the cost or capability case for moving is now substantial. If you buy meeting software rather than build it, the right response is to do nothing and expect your existing vendor's transcripts to quietly improve over the next few quarters.

Sources

Related tool reviews

Questions or corrections? Email Pick Right. Want the full list? See all news.