Alibaba just dropped speech-to-text through the price floor — and changed the meter on the way down
TL;DR: Alibaba shipped Qwen-Audio-3.1 at the Apsara Conference on 22 September 2026 — five models covering recognition, synthesis, real-time conversation and audio creation — with cuts of ~70% on TTS, ~85% on Realtime, up to 95% on ASR. Worked against Qwen’s documented 25 audio tokens per second, ASR-Flash lands near $0.016 per audio-hour. The cheap tier of speech-to-text has sat at $0.18–$0.30 for years. This is roughly a tenth of it — the first time the floor has genuinely moved rather than been joined. The catches: the meter changed from audio-seconds to tokens on the same day, so the percentages compare two different units; part of your bill now scales with how densely people talk; two of the five models have no API yet; the rate card quoted is the Beijing node, and Alibaba’s international pricing page does not list the family at all.
What shipped
On 22 September 2026, at Alibaba Cloud’s Apsara Conference, the Qwen team replaced its audio line with five models under one version number.
Three are upgrades. ASR widens dialect coverage to 30 languages and 16 Chinese dialects and adds native transcript polishing — filler and repetition removal at the model layer rather than in your post-processing — with first-character response around 160 milliseconds. TTS adds cross-lingual timbre transfer, holding one voice’s identity across Mandarin, Cantonese, English and Japanese, and replaces version 3.0’s 86 inline delivery tags with natural-language instructions for emotion, speed and style. Realtime goes full-duplex: speaking and listening at once, interruptible at any point.
Two are new. ASR-Next does multi-speaker recognition with speaker labels, timestamps and aligned transcripts, and extends past speech into emotion, ambient sound and machine-noise understanding — sound captioning, event localisation, audio question-answering. TTS-Next runs a combined language-model-and-diffusion framework that generates dialogue, sound effects and background audio in a single pass, aimed at audiobooks, podcasts, games and ads.
Alongside them Alibaba cut simultaneous-interpretation latency in Qwen3.8-LiveTranslate by nearly 20%, from 2.8 to 2.3 seconds.
That is a complete audio stack, announced competently. It is also not the reason this release matters.
The arithmetic nobody ran
The coverage has settled on three percentages: TTS down about 70%, Realtime about 85%, ASR up to 95%. Percentages are not prices. Here is the conversion.
Alibaba’s Beijing-node rate for Qwen-Audio-3.1-ASR-Flash is ¥0.8 per million input tokens and ¥2.7 per million output tokens. At the 23 September 2026 rate of 6.71 yuan to the dollar, that is $0.119 and $0.403.
Qwen’s documented ASR tokenisation is 25 audio tokens per second. An audio-hour is therefore 3,600 × 25 = 90,000 input tokens, costing $0.0107. An hour of speech transcribes to roughly 9,000 words — call it 12,000 output tokens — adding $0.0048.
Roughly $0.016 per audio-hour.
For scale, our analysis of Gemini 3.5 Transcribe in August mapped the cheap tier of this market, and it had barely moved in years:
| Vendor | Loaded cost per audio-hour |
|---|---|
| Qwen-Audio-3.1-ASR-Flash | ~$0.016 |
| Meta Muse Voice Transcribe | $0.18 |
| AssemblyAI | $0.23 |
| Deepgram Nova-3 | $0.28 |
| Gemini 3.5 Transcribe | $0.30 |
| OpenAI Whisper API (no diarization) | $0.36 |
The thesis of that August piece was that Google had walked down to a floor the specialists set years ago, not reset it. That reading was right then and it is what makes this different now. Qwen-Audio-3.1-ASR is roughly eleven times below Meta’s $0.18 and fourteen times below Gemini’s $0.30. This is the first move in this category that is a break rather than a convergence.
The meter moved on the same day
Now the part that should slow you down.
The predecessor, qwen3-asr-flash, billed by audio duration — $0.000035 per second on the international deployment, or $0.126 per audio-hour. Version 3.1 bills per token. On the synthesis side, Qwen-Audio-3.0-TTS billed per character; 3.1 bills per token.
Against that $0.126 duration price, our $0.016 works out to an 88% cut — real, large, and not the 95% on the banner. The gap is not necessarily spin; the 95% may describe a different tier or the domestic rate card. But it cannot be checked, because the unit changed underneath the comparison and Alibaba did not republish the audio-token conversion rate alongside the new prices. The 25-tokens-per-second figure is documented for Qwen ASR generally, not restated for 3.1.
This is the second time this month a voice vendor has repriced by changing what it counts. When GPT-Live-1 went to a flat per-minute meter on 10 September, the headline was a price cut and the substance was that silence stopped being free — hold queues and thinking pauses became billable. Qwen has run the same play in the opposite direction, and the consequence is the mirror image.
Duration and character billing are quotes. An hour is an hour; a million characters is a million characters. Token billing is an estimate, split in two. The audio input side stays fixed, because audio tokenises at a constant rate no matter what is on the recording. The output side moves with content: a dense four-person technical panel generates far more transcript per audio-hour than a one-on-one with long pauses. On the numbers above, that variable component is about 30% of the blended cost.
That is precisely the structural criticism we levelled at Gemini 3.5 Transcribe’s token billing in August — now adopted by the vendor undercutting it. At a tenth of the price the variance matters far less in absolute terms. It still means your forecast is a model, not a quote.
There is one encouraging cross-check. Run the same arithmetic on Qwen-Audio-3.1-TTS-Flash at ¥1.5 input and ¥12 output per million tokens and you get roughly $3.30 per million synthesised characters, against about $15 for the 3.0 Flash tier it replaces — a 78% cut, close enough to Alibaba’s claimed ~70% that the tokenisation assumption appears sound. Two independent derivations landing near the vendor’s own figures is the best verification available without a published conversion table.
What you cannot buy yet
Four things stand between this rate card and a migration.
Two of the five models have no API. ASR-Next and TTS-Next — the multi-speaker recognition and the single-pass audio-scene generation, which are the genuinely novel entries — were still listed as pending at announcement. They are demos.
The rate card is not in your region. Every yuan figure above is the Beijing node. The international pay-as-you-go path runs through the Singapore-anchored ap-southeast-1 endpoint in USD at different rates — realtime transcription there moved to $0.93 per million input and $0.70 per million output tokens. At the time of writing, Alibaba Cloud’s international Model Studio pricing documentation did not list the Qwen-Audio-3.1 family at all. You cannot sign a contract against a page that does not exist.
Beijing is a data-residency answer, not a location detail. Call recordings and meeting audio are among the most sensitive data most organisations hold. For an EU or US buyer this is a GDPR and procurement question that precedes the cost question entirely — the same split we saw when Grok 4.6 landed on Bedrock with a data-residency premium, where the cheap endpoint and the compliant endpoint were not the same endpoint.
Quality is unmeasured. Qwen-Audio-3.1 has no independent score. The verified 1,259 Elo on the Artificial Analysis Provider Voice Arena belongs to the 3.0 Plus tier — the model this release supersedes. A price an order of magnitude below the market is only interesting if the output is usable, and right now the only evidence for that is the vendor’s own.
The synthesis side, and why price is not the argument
For text-to-speech the gap is wider still and the case is weaker.
At roughly $3.30 per million characters, Qwen sits against ElevenLabs Eleven v3 at about $100 per million characters as tracked by Artificial Analysis, with Flash near $50. A 20-million-character monthly narration workload is roughly $2,000 on Eleven v3 and roughly $66 on Qwen.
A 30x gap will get a procurement meeting. It should not win it on its own, because characters are not what you are buying.
The day after Qwen’s announcement, Google shipped Gemini 3.8 Flash TTS with consent-gated 30-second voice replication, SynthID and C2PA provenance marking — and blocked voice replication in the EU outright rather than test Article 50. Qwen’s announcement says nothing about provenance marking at all. For anyone publishing synthetic audio into Europe, that silence is a compliance exposure that no per-character saving offsets, and it is the same gap that made Lyria 3.5’s SynthID marking the deciding factor in music generation rather than output quality.
There is also a licence-posture question. Qwen has changed terms under shipped products before — Qwen-Image 2.1 moved from Apache 2.0 to research-only on 21 September, a month after launch. That is not a reason to avoid Qwen; it is a reason to keep your audio layer behind an abstraction you can repoint.
What to actually do
If you call a speech-to-text API directly, this is worth a bounded pilot — but on the international endpoint, with your own densest audio, measuring blended cost per audio-hour rather than trusting the per-token rate. The number to beat is whatever you pay now; if you are on a legacy enterprise stack like Google Cloud STT or AWS Transcribe, the case was already overwhelming before this release and is now absurd.
If you are on a specialist — Deepgram, AssemblyAI — price is now genuinely a reason to look, which it was not in August. It is still not a reason to move until residency, provenance and an international rate card exist. Use it in your renewal.
If you synthesise voice for publication in the EU, stay where provenance marking is documented. The gap is real and it is not the one on the invoice.
If you buy meeting or podcast software rather than build it — Otter.ai, Descript, and the rest of the audio tooling shortlist — do nothing. This is an API-layer release. Your vendor’s input costs just fell by an order of magnitude; you will experience that as slightly better transcripts and, eventually, as pricing pressure someone else has to apply.
The broader pattern is the one we have been tracking since the price war flipped in August: Chinese labs set the floor, Western labs walk down to it, and the floor is increasingly set at a level that looks less like a margin and more like a customer-acquisition budget. What is new here is that for the first time in speech, the floor did not hold. Developers building on audio should plan for a market where recognition is effectively free within two years — and where the thing you are actually paying for is the contract, the region and the provenance trail, not the tokens.
Prices converted at 6.71 CNY/USD as of 23 September 2026 and will move. Per-audio-hour figures are derived from Qwen’s documented 25-audio-tokens-per-second rate and an assumed 9,000 words per hour of speech; Alibaba has not published a conversion table for Qwen-Audio-3.1. Verify against your own audio before committing budget.
Update, 25 September 2026 — the other way to move a price is to delete it
This article’s thesis is that the meter change is the story, because a percentage quoted across a changed unit cannot be checked against your old invoice. Two days later Google demonstrated the limit case.
On 24 September, Gemini 3.8 Live with Live Avatar reached general availability with real-time generated video attached to Google’s speech-to-speech models — and no published rate for the video stream. The conversation bills on the existing Live API audio meter ($3 and $12 per million tokens, roughly $0.011 a conversational minute at Google’s own 25-audio-tokens-per-second rate — the same tokenisation constant used in the Qwen arithmetic above). The avatar itself is presented as a feature, not a product.
For the comparison shopper this is harder to handle than Alibaba’s 95%. A changed unit at least leaves you a number to convert; a missing unit leaves nothing to forecast against, and nothing preventing a rate appearing later. HeyGen charges $0.10 a minute in LITE mode for the equivalent layer, which is the market’s current opinion of what it is worth. That figure, not zero, is the number to put in a budget.
Frequently asked questions
What exactly did Alibaba ship, and can I use all of it today?
Five models under the Qwen-Audio-3.1 name, announced at the Apsara Conference on 22 September 2026. Three are upgrades to existing lines: ASR adds stronger dialect coverage — 30 languages and 16 Chinese dialects — plus native transcript polishing that strips fillers and repetitions, with first-character response around 160 milliseconds. TTS adds cross-lingual timbre transfer, so one voice keeps its identity across Mandarin, Cantonese, English and Japanese, and swaps 3.0's 86 inline delivery tags for natural-language instructions on emotion, speed and style. Realtime adds full-duplex conversation — speaking and listening simultaneously with interruption at any point. Two are new: ASR-Next does multi-speaker recognition with speaker labels and timestamps plus emotion, ambient and machine-sound understanding; TTS-Next uses a combined language-model-and-diffusion framework to generate dialogue, sound effects and background audio in a single pass. The availability caveat matters: the ASR-Next and TTS-Next APIs were still listed as pending at announcement. Two of the five headline models are demos, not purchasable endpoints.
Is speech recognition really 95% cheaper, and what does it actually cost per hour?
The direction is real, the specific number is not verifiable, and the honest figure is closer to 88% than 95% on the tier we can price. Alibaba's Beijing-node rate for Qwen-Audio-3.1-ASR-Flash is ¥0.8 per million input tokens and ¥2.7 per million output tokens — about $0.119 and $0.403 at the 6.71 yuan-to-dollar rate of 23 September 2026. Qwen's documented ASR tokenisation is 25 audio tokens per second, so an audio-hour is 90,000 input tokens, costing roughly $0.0107. An hour of speech transcribes to around 9,000 words, call it 12,000 output tokens, adding about $0.0048. Total: roughly $0.016 per audio-hour. The predecessor, qwen3-asr-flash, billed by duration at $0.000035 per second on the international deployment — $0.126 per audio-hour. That is an 88% reduction. The 95% headline likely applies to a different tier or the China-node rate card, and Alibaba has not shown its working. Either way the conclusion holds: this is roughly a tenth of the cheap tier's established price.
How does this compare to Deepgram, AssemblyAI, Google and Meta?
It is an order of magnitude below all of them, which is why it matters. The cheap tier of speech-to-text has been remarkably stable: AssemblyAI at about $0.23 per hour loaded with diarization, Deepgram Nova-3 at about $0.28, Meta's Muse Voice Transcribe at $0.18, Gemini 3.5 Transcribe at about $0.30, OpenAI's Whisper API at $0.36 with no diarization at all. When Google launched Gemini 3.5 Transcribe in August, the accurate reading was that Google had walked down to a floor the specialists set years ago, not reset it. Qwen-Audio-3.1-ASR at roughly $0.016 an hour is a genuine break — about 11 times below Meta's $0.18 and 14 times below Gemini's $0.30. Whether it is a floor you can build on is a separate question from whether it is a floor: the specialists sell contractual per-minute quotes, uptime commitments and regional endpoints, and a per-token rate on a Beijing node is not the same instrument even when the arithmetic is a tenth of the price.
What is the catch with the billing-unit change?
Two things, one structural and one about verification. Structurally, ASR moved from per-audio-second to per-token and TTS moved from per-character to per-token. Duration and character billing are quotes: an hour is an hour and a million characters is a million characters, so silence costs the same as speech and you can forecast exactly. Token billing splits your bill in two. The audio input side stays effectively fixed, because audio tokenises at a constant 25 tokens per second regardless of content. The output side does not — a dense four-person technical panel produces far more transcript tokens per audio-hour than a sparse one-on-one with long pauses. On our worked ASR numbers the variable half is about 30% of the blended cost. That is the same structural weakness we flagged in Gemini 3.5 Transcribe's token billing, now adopted by the vendor undercutting it. On verification: because the unit changed on the same day as the cut, the percentages compare two different meters, and Alibaba did not republish the audio-token conversion rate alongside the new prices. The 25-tokens-per-second figure is documented for Qwen ASR generally, not restated for 3.1. Run your own densest audio through it before extrapolating.
Should I actually move my transcription or voice workload to Qwen?
For most Western buyers, not yet, and the blockers are not about price. First, data residency. The yuan rate card is the Beijing node; that is China-hosted processing, which for most EU and US organisations handling call recordings or meeting audio is a GDPR and procurement conversation before it is a cost conversation. The international Singapore endpoint bills in USD at different rates — realtime transcription there moved to $0.93 per million input and $0.70 per million output tokens — and at the time of writing Alibaba's international Model Studio pricing page did not list the Qwen-Audio-3.1 family at all. You cannot sign against a rate card that is not published in your region. Second, quality is unmeasured: 3.1 has no independent score yet, and the 3.0 Plus tier's verified 1,259 Elo on the Artificial Analysis Voice Arena tells you about a model this release replaces. Third, two of the five models have no API. Fourth, Qwen's licence posture has moved before — Qwen-Image 2.1 went from Apache to research-only in September. The reasonable move is a bounded pilot on the international endpoint with your own audio, not a migration. If you buy meeting software rather than call an API, do nothing: your vendor's margin improves and your transcripts get quietly better.
Does this change anything for text-to-speech buyers on ElevenLabs?
It changes the negotiation more than the shortlist. Applying the same tokenisation arithmetic to Qwen-Audio-3.1-TTS-Flash at ¥1.5 per million input and ¥12 per million output tokens gives roughly $3.30 per million synthesised characters, against about $15 for the 3.0 Flash tier it replaces — a 78% cut that lands close enough to Alibaba's claimed 70% to serve as a cross-check on the whole exercise. ElevenLabs Eleven v3 runs near $100 per million characters as tracked by Artificial Analysis, with Flash around $50. On a 20-million-character-a-month narration workload that is roughly $2,000 monthly on Eleven v3 against roughly $66 on Qwen. But characters are not the product — voice quality, voice cloning consent controls, provenance marking and the ability to point at a contract are. Google's Gemini 3.8 Flash TTS shipped the day after Qwen with consent-gated voice replication and SynthID plus C2PA provenance, and blocked replication in the EU entirely rather than risk Article 50. Qwen's announcement says nothing about provenance marking. If you are producing published audio in Europe, that silence is a larger problem than a 30x price gap is an opportunity.
Sources
- Qwen (@Alibaba_Qwen) — "Meet Qwen-Audio-3.1": five-model stack announcement and price cuts
- Qwen Developers (@QwenDevs) — full audio workflow: ASR, TTS, Realtime upgraded; ASR-Next and TTS-Next join
- Alibaba — Unveils Roadmap on Full-Stack AI Strategy from Chips, Cloud Infrastructure, Models to Agents (Apsara Conference, 22 September 2026)
- The Decoder — Alibaba launches Qwen Audio 3.1 with five new models and slashes AI audio prices by up to 95 percent
- 36Kr — Alibaba Unveils 5 New Qwen AI Models Simultaneously with Up to 95% Price Cut (per-model yuan rate card)
- Pandaily — Alibaba Upgrades Qwen Audio Suite: LiveTranslate Latency Cut and Qwen-Audio-3.1 ASR/TTS/Realtime
- AlphaSignal — Alibaba's Qwen-Audio 3.1 Slashes Voice API Prices by up to 95% (model IDs; ASR-Next and TTS-Next APIs pending)
- OrcaRouter — Qwen-Audio-3.1-TTS vs Qwen-Audio-3.0-TTS: What Changed (per-character to per-token billing shift)
- OrcaRouter — Qwen-Audio-3.1-TTS vs ElevenLabs: A 70% Price Cut Arrives
- Alibaba Cloud Model Studio — model pricing documentation (international rate card)
- Portkey AI models registry — qwen3-asr-flash billed by audio duration at $0.000035/second (international deployment)
Related tool reviews
Questions or corrections? Email Pick Right. Want the full list? See all news.