Google releases Gemini 3.8 Flash TTS

24-09-2026

Google released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS on 23 September 2026. Flash-Lite costs $6 per million audio output tokens until the end of the year, then doubles. Both carry over 2,000 voices in more than 100 languages and let you design a new voice from a written description.

Written by:

Jorick van Weelie

Marketing Lead at DataNorth | AI Enthusiast & Tech Storyteller

google releases gemini 3.8 flash tts with 2,000 voices in over 100 languages
Sign up for our Newsletter

Published: 24 September 2026

Google released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS on 23 September 2026. Flash-Lite TTS costs $6 per million audio output tokens until the end of this year, then doubles to $12. Both models ship with more than 2,000 ready-made voices across over 100 languages, and you can invent a new voice by describing it in a sentence.

What can Gemini 3.8 Flash TTS do?

It turns text into speech, and it lets you direct the performance. You write a stage direction next to a line, such as a whisper or a rushed delivery, and the model follows it line by line. That is the part the old Gemini TTS models could not do.

The voice library is the second change. Google went from 30 fixed voices to over 2,000, and generative voice design lets you write a description instead of picking from a list. You can also clone a voice from a 30 second sample, though you have to record the speaker giving verbal consent first.

The two models split the work. Flash TTS is the quality tier, aimed at narration, character voices and multi-speaker dialogue. Flash-Lite TTS is the volume tier, aimed at dubbing and voice agents where you are generating audio all day.

What you get with both:

  • Model IDs: gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts
  • Over 100 languages and dialects, including Quebec French and Scots English
  • SynthID watermarking on every audio file the models produce
  • Long-form generation that holds quality across hours of continuous audio
  • Available now in the Gemini API and Google AI Studio, with Gemini Enterprise coming later
  • API only, so there are no weights to self-host

Gemini 3.8 Flash TTS pricing and benchmarks

Take this from the table: you pay the same to send text in, so the real choice is what an hour of finished audio costs you.

MeasurementGemini 3.8 Flash TTSGemini 3.8 Flash-Lite TTS
Text input, per million tokens$0.50$0.50
Audio output to 31 December 2026, per million tokens$9.00$6.00
Audio output from 1 January 2027, per million tokens$18.00$12.00
Hume AI Overall Quality Index (how good it sounds)1st2nd
Hume AI Voice Design Benchmark (matching a described voice)71.4, 1st overallnot published
Accent modelling, same benchmark60.8, 1stnot published

The prices come from Google’s own pricing page, not from the launch post. The rankings are Hume AI’s, a third party, measured on its Voice Design Benchmark and its Overall Quality Index. Google published no evaluation of its own for either model, which is unusual and, on balance, a point in its favour: an outside scoreboard is harder to tune towards.

What Google is not saying about the new TTS models

The price doubling on 1 January 2027 appears only in the pricing table. The launch post does not mention it. If you are building a business case on $6 per million audio tokens, you have about three months at that rate.

There is also no latency figure anywhere. For narration that does not matter. For a voice agent it is the number that decides whether the product feels alive, and Google has published nothing on time to first audio byte. Billing is per token rather than per second or per character, so you cannot work out the cost of a minute of speech from the published figures either. You have to measure it.

Voice replication is blocked in Illinois, Texas, the EEA, the UK, Switzerland and India. That rules it out for most European teams.

What this means

Worth testing now if you generate speech at volume. A dubbing team or a support platform running voice agents should put Flash-Lite TTS side by side with whatever they use today this week, because $6 per million audio tokens undercuts most of the specialist voice vendors, and 100 languages from one endpoint removes a pile of integration work. Measure latency yourself on your own traffic, since Google has not published it, and price the 2027 rate into anything with a twelve month horizon.

Worth watching, not switching, if you are a European team whose main interest was voice cloning. That feature is unavailable to you, and what remains is a very good stock voice library rather than the thing you wanted. The wider point is that Google is now competing on price in a market that specialist vendors built, and it has chosen to be judged on someone else’s benchmark rather than its own. Both of those are pressure on ElevenLabs and on OpenAI’s audio models, and neither has answered yet.

For more information, visit the official announcement of Gemini 3.8 Flash TTS on the Google blog.

Add DataNorth AI to your Google favorites