Google launches Gemini 3.8 Live and Extended Thinking

16-09-2026

Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on 15 September 2026. Extended Thinking tops the Artificial Analysis Speech to Speech Index at 82.6, just ahead of OpenAI's GPT-Live-1, and costs $3.50 per hour of input audio against $5.83. The cheaper Gemini 3.8 Live runs at $0.84 per hour but completes only 30.1% of agentic voice tasks. Both are live in the Gemini API today.

Written by:

Diederik Knol

Online Marketeer at DataNorth | Passionate about AI

Google launches Gemini 3.8 Live and Extended Thinking
Sign up for our Newsletter

16 September 2026

Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on 15 September 2026, two voice models that listen and speak at the same time. Gemini 3.8 Live Extended Thinking scores 82.6 on the Artificial Analysis Speech to Speech Index, the highest of any voice model that firm has tested, and runs at $3.50 per hour of input audio against $5.83 for OpenAI’s GPT-Live-1. Both models are available today in the Gemini API and Google AI Studio.

What can Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking do?

These are speech-to-speech models: audio goes in, audio comes out, with no separate transcription or text-to-speech step in between. Cutting out those steps is what lets the model deal with an interruption mid-sentence instead of waiting for you to finish talking.

The two models split the work. Gemini 3.8 Live is the fast, cheap one, aimed at high call volumes. Gemini 3.8 Live Extended Thinking reasons and speaks at the same time, so it can say “let me check that” and then narrate its progress while a multi-step task runs in the background.

Both take audio, images, video and text, with a 128,000-token context window and 64,000 tokens of output. Gemini 3.8 Live reads visual input in near real time and detects and switches between 97 languages mid-conversation. All audio output carries a SynthID watermark. Google’s model card states that both models are built on Gemini 3 Pro rather than on a Flash model.

Gemini 3.8 Live benchmarks and how it compares to GPT-Live-1

OpenAI shipped GPT-Live-1 in its API five days earlier, so the two land in the same week. The table shows why the choice is close on quality and not close on price.

What is measuredGemini 3.8 Live Extended ThinkingOpenAI GPT-Live-1 (Astra, medium)
Speech to Speech Index (overall quality)82.681.5
Big Bench Audio (reasoning from speech)98%90%
Tau-Voice (customer service tasks completed)68.6%67.9%
Full Duplex Bench (interruptions and pauses)91.9%94.9%
Time to first audio1.35 seconds1.34 seconds
Cost per hour of input audio$3.50$5.83

These are Artificial Analysis measurements, not Google’s, although Google quotes the same Tau-Voice and Big Bench Audio figures in its launch post. Artificial Analysis notes that the GPT-Live-1 row rests on a single trial where most models get three, so treat the OpenAI column as the noisier of the two.

What Gemini 3.8 Live costs

Google’s price list puts Gemini 3.8 Live, Gemini 3.8 Live Extended Thinking and the older Gemini 3.1 Flash Live Preview on a single line, so the two new models cost the same per token:

  • Audio input: $3.00 per million tokens, or $0.005 per minute
  • Audio output: $12.00 per million tokens, or $0.018 per minute
  • Text input $0.75 and text output $4.50 per million tokens
  • Image and video input: $1.00 per million tokens, or $0.002 per minute
  • A free tier is available through Google AI Studio

Identical list prices do not mean identical bills. Extended Thinking generates reasoning tokens, and those are charged at the output rate. On the same fixed task set, Artificial Analysis measured $3.50 per hour of input audio for Extended Thinking against $0.84 for plain Gemini 3.8 Live: the same price card, roughly four times the invoice.

What Google is not saying

Three things sit outside the launch post. The first is the margin. A score of 82.6 beats GPT-Live-1’s 81.5 by 1.1 points on a composite index. Google calls it the number one overall spot and leaves the gap unstated.

The second is the weaker half of the pair. Gemini 3.8 Live completes 30.1% of Tau-Voice customer service tasks against 68.6% for Extended Thinking. Google prints the 68.6% and says only that Gemini 3.8 Live took second place in the Speech Agent Arena. Arena position measures which model people preferred talking to. It does not measure whether the task got finished. If you are building an agent that has to change a flight or refund an order, 30.1% is the number that decides it. That second place is also behind Google’s own Gemini 3.1 Flash Live Minimal, not behind a rival.

The third is the safety paperwork. The model card says Google evaluated Gemini 3.7 Flash under its Frontier Safety Framework and carried the conclusion across, on the grounds that neither new model shows meaningful new capabilities. No fresh frontier-safety evaluation was run on either release. The same card gives a knowledge cutoff of January 2025, twenty months before launch.

What this means

If you already run a voice agent on GPT-Live-1 or GPT-Realtime, Gemini 3.8 Live Extended Thinking is worth testing this week, and the reason is cost rather than quality. A 1.1-point index lead sits inside the noise of any benchmark. A 40% lower cost per hour of audio does not. For a contact centre team handling thousands of calls a day, that difference lands on the monthly invoice immediately, and the switching cost is a few days of harness work.

Treat the two models as separate products rather than as a fast tier and a slow tier of one thing. Gemini 3.8 Live at $0.84 per hour is the cheapest credible voice model on the board and it is pleasant to talk to, so it fits triage, routing, opening-hours questions and anything where a person picks up the hard case. It does not fit unsupervised task completion, and the 30.1% Tau-Voice score says so plainly. Extended Thinking is the one to point at real workflows. Budget for it as a different model with a different bill, because that is what it is.

For more information, visit the official announcement of Gemini 3.8 Live on the Google blog.

Add DataNorth AI to your Google favorites