Google has launched Gemini 3.8 Live, a speech-to-speech model that holds a spoken conversation and runs tools while it talks. It costs $0.005 a minute for audio in and $0.018 a minute for audio out, and covers 97 languages. Google says a companion version, Extended Thinking, ranks first on Artificial Analysis’s voice quality index.
What Gemini 3.8 Live Does
Gemini 3.8 Live is a native speech-to-speech model, which means it takes audio in and produces audio out without converting to text in between. That is what lets it react to an interruption, a change of tone or a switch of language mid-sentence instead of restarting the exchange.
Google, the Mountain View company that develops the Gemini model family, lists four behaviours that separate it from earlier voice models:
- Background tool calls: The model can query an API or run a function while it keeps speaking, rather than going silent during the lookup.
- Visual grounding: It can read what a camera is pointed at in near real time and answer about it.
- Language switching: It detects the language being spoken across 97 languages and changes mid-conversation without being told to.
- Alphanumeric precision: Google singles out accuracy on strings such as reference numbers and codes, the failure case that makes voice agents unusable for account queries.
Every audio output carries Google’s SynthID watermark, which the company describes as an “imperceptible watermark… woven directly into the audio output”.
What an Hour of Conversation Costs
The headline figure circulating is $1.38 an hour, and it is worth knowing how that number is built, because it is a worst case rather than a typical bill.
Google’s developer announcement prices the model at $0.005 per minute of audio input and $0.018 per minute of audio output. Sixty minutes of each, billed together, comes to $1.38. Real conversations do not have both sides talking for the full hour, so a call where the model speaks for a third of the time costs closer to $0.66.
Output is 3.6 times the price of input, which is the practical design constraint: a voice agent that listens patiently is cheap, and one that monologues is not.
How the Two Models Differ
| Gemini 3.8 Live | Gemini 3.8 Live Extended Thinking | |
|---|---|---|
| Built for | Fast, high-volume voice agents | Multi-step reasoning inside a live call |
| Audio input | $0.005 per minute | $0.005 per minute |
| Audio output | $0.018 per minute | $0.018 per minute |
| Extra charges | None published | Reasoning tokens, plus video and document inputs, per SiliconANGLE |
| Languages | 97 | 97 |
The two share a per-minute audio price. The difference is that Extended Thinking bills for the reasoning it does on top, a detail reported by SiliconANGLE rather than stated in Google’s consumer announcement, so anyone budgeting for it should expect the audio rate to be a floor.
The Benchmark Scores Google Published
Google put Extended Thinking at the top of the main voice leaderboard and published the supporting numbers.
| Benchmark | Result |
|---|---|
| Artificial Analysis Speech to Speech Quality Index | 82.6, first overall (Extended Thinking) |
| Big Bench Audio | 97.7% |
| τ-Voice, agentic task completion | 68.6% |
| τ-Voice-banking (Sierra) | 35.1% |
| Artificial Analysis Speech Agent Arena | Second place (Gemini 3.8 Live) |
The 82.6 score and the first-place claim come from Google’s own announcement, and SiliconANGLE reported the same figures independently. The banking result is the one to note: 35.1% on a benchmark of real customer-service tasks is a long way from solved, and it is Google’s own published number.
Which Google Apps Use It, and on Which Plan
This is the part most coverage skipped. The models are not only a developer product; they are already behind features in Google apps, and the access rules differ by app.
- Gemini Live: Extended Thinking began rolling out to the Gemini app’s live mode on 15 September 2026.
- Gmail and Keep: Extended Thinking is available to all subscribers, according to Google’s announcement.
- Google Workspace: Restricted to Google AI Pro and Ultra subscribers.
- Search Live: Runs on the standard Gemini 3.8 Live model and is public.
- Gemini Enterprise: Private preview only, with Gemini Enterprise for Customer Experience listed as coming.
As of 17 September 2026, both models are available to developers through the Gemini API and Google AI Studio, and through the partner platforms Vercel, Agora, LiveKit, Pipecat, Fishjam, LangChain and Vision Agents. Google has not published a rollout schedule by country for the consumer features.
For readers tracking where Gemini is turning up next, we covered the Gemini app arriving on Windows earlier this month, and the retirement of Google Assistant in Gemini’s favour.
Gemini 3.5 Transcribe Shipped the Same Day
A third model launched alongside the two voice models and received almost no attention. Gemini 3.5 Transcribe is speech-to-text rather than speech-to-speech, covering 85 languages.
Google reports a word error rate of 4.0% in streaming mode and 2.6% when transcribing a completed file, and says the Interactions API accepts files of up to one hour. It also handles code-switching automatically and allows custom vocabulary biasing, which matters for names and jargon a general model mishears.
What Google Has Not Disclosed
Three gaps stand out. Google has not published the token rate that Extended Thinking charges for reasoning, so the true cost of that variant cannot be calculated from the announcement alone.
It has also given no date for general availability of the enterprise tiers beyond “private preview”, and no country-by-country schedule for the consumer features. Nothing in the announcement compares latency with rival voice models in milliseconds, which is the figure developers choosing between providers would most want.
Gemini 3.8 Live FAQ
How Much Does Gemini 3.8 Live Cost?
$0.005 per minute of audio input and $0.018 per minute of audio output. An hour of continuous audio in both directions works out to $1.38, and a normal call costs less because both sides are not talking at once.
Is Gemini 3.8 Live Available to Ordinary Users or Only Developers?
Both. Developers can use it through the Gemini API and Google AI Studio, and the Extended Thinking version is rolling out inside Gemini Live, Gmail, Keep and Google Workspace, with Search Live using the standard model.
Do I Need a Paid Google Plan?
It depends on the app. Google says Gmail and Keep get Extended Thinking for all subscribers, while the Workspace features require a Google AI Pro or Ultra subscription.
How Many Languages Does Gemini 3.8 Live Support?
97, with automatic detection and switching during a conversation. The separate Gemini 3.5 Transcribe model covers 85 languages for speech-to-text.
Can You Tell Whether Audio Came From Gemini 3.8 Live?
Google says all generated audio carries a SynthID watermark embedded in the output itself. Detecting it requires Google’s tooling rather than listening for it.




