Published: 27 August 2026
Google launched Gemini 3.5 Transcribe on 26 August 2026. It turns speech into text with a 2.6 percent word error rate, so it gets roughly 97 of every 100 words right. It replaces Chirp 3 and is now in public preview for developers.
What can Gemini 3.5 Transcribe do?
Most speech tools hand you a raw transcript, filler words and all. Gemini 3.5 Transcribe cleans as it goes. It drops the ums and ahs, adds punctuation and formatting, and follows corrections. Say let us meet Tuesday, no, Wednesday, and it writes Wednesday. It can also hand heavier jobs, such as generating an image, to other Gemini models.
It picks up the language on its own across more than 85 of them. In recordings it labels up to three speakers with timestamps, though more than three is still experimental. You can also give it a list of your own jargon and product names so it spells them correctly.
Gemini 3.5 Transcribe accuracy and availability
Google measures the model against Chirp 3, its 2025 transcription model, using figures from Artificial Analysis.
| What is measured | Gemini 3.5 Transcribe |
|---|---|
| Word error rate, recorded audio | 2.6 percent |
| Word error rate, live streaming | 4.0 percent |
| FLEURS multilingual test, recorded | 5.04 percent |
| FLEURS multilingual test, streaming | 5.50 percent |
| Speed to final transcript | 70 percent faster than Chirp 3 |
| Languages detected automatically | more than 85 |
| Speakers labelled in a recording | up to 3 |
These figures come from Artificial Analysis rather than from Google’s own lab, which is worth noting, because vendors usually grade their own homework. The gap between 2.6 and 4.0 percent also matters: live captions will be visibly worse than a processed recording. Google has not published a price, so you cannot build a business case yet.
Where it is already running
This is not a limited trial. The model already powers dictation in Gboard on Android, voice control in the Gemini app for Mac, and the microphone in Google Antigravity. Chrome is next, which will let you dictate into any web page.
That breadth is the real signal here. Google is switching its whole product line off Chirp 3 in one move. When a company retires its own model across Search, Docs, Gmail and Keep at the same time, it is confident the replacement holds up.
Google also published GlucoFM
On the same day, Google Research and UNSW Sydney published GlucoFM, a small model that reads continuous glucose data. It splits a blood sugar trace into a slow underlying pattern and the short spikes caused by meals or stress. At 0.72 million parameters it beat far larger models on the tests reported.
Do not get excited yet. There is no model to download, no regulatory approval, and every test looked backwards at old data. Google says so itself. What is useful is the recipe, which trains on a single GPU.
What this means
Google Transcribe 3.5 is worth testing now if you build anything that listens: meeting notes, call centre transcripts, voice agents, or dictation inside your own app. A 2.6 percent error rate checked by an outside party is a stronger claim than most transcription vendors can make. Start with recorded audio, where accuracy is best, before you promise anything live.
Two things will slow you down. There is no published price, so finance cannot approve anything yet. And speaker labelling stops at three voices, which rules out most real meetings. If you record five people around a table, this model will not sort them out.
For more information, visit the official announcement of Gemini 3.5 Transcribe on the Google blog.