Opens in a new tab

Microsoft releases MAI-Transcribe-2-Streaming

02-10-2026

Microsoft released MAI-Transcribe-2-Streaming on 1 October 2026, its first real-time transcription model, with MAI-Voice-2.1 and MAI-Voice-2.1-Flash. Microsoft reports a 2.5% word error rate and a final transcript 0.13 seconds after speech ends, across 60 languages with automatic detection.

Written by:

Jorick van Weelie

Marketing Lead at DataNorth | AI Enthusiast & Tech Storyteller

microsoft releases mai transcribe 2 streaming with a 2.5% word error rate
Sign up for our Newsletter

Microsoft released MAI-Transcribe-2-Streaming on 1 October 2026, its first real-time transcription model, alongside two speech models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash. Microsoft reports a 2.5% word error rate with the final transcript ready 0.13 seconds after someone stops talking. Streaming costs $0.54 per hour of audio as an introductory rate through the end of 2026.

What is new in MAI-Transcribe-2-Streaming?

Microsoft already had MAI-Transcribe-2, a batch model that takes a finished recording and returns text. Streaming is a different problem. The model has to show words while the person is still speaking, then quietly correct them as more audio arrives. Microsoft says the first words appear just over 100 milliseconds after audio starts flowing, roughly twice as fast as its nearest rivals.

It covers 60 languages with automatic language detection, so you do not have to tell it which language is coming. That matters for a support line that takes calls in a dozen countries on one number.

The two voice models go the other way, turning text into speech. MAI-Voice-2.1 handles 23 languages across 26 locales and keeps one recognisable voice across all of them, so a bilingual assistant does not change character mid sentence. MAI-Voice-2.1-Flash is the cheap, fast version: 150 milliseconds end to end to begin producing 45 seconds of audio, with inference Microsoft measures as 55% faster than the standard model. Both support voice cloning with consent checks built in.

MAI-Transcribe-2-Streaming benchmarks and pricing

The accuracy gap is small and the latency gap is large, which is the opposite of how the launch reads.

AspectMAI-Transcribe-2-StreamingCompetitor
Word error rate (mistakes per 100 words)2.5%Grok Voice Transcribe 2.0: 2.73%
Final transcript ready after speech ends0.13 secondsGrok Voice Transcribe 2.0: 0.49 s, Muse Voice Transcribe: 0.16 s
Languages covered60, with auto detectionMAI-Voice-2.1: 23 languages, 26 locales
Streaming price per 1,000 minutes$9.00 introductoryElevenLabs and Deepgram Flux: about $6.50
Price against Microsoft’s own batch model$0.54 per hourMAI-Transcribe-2 batch: $0.10 per hour
Speech output per 1 million charactersMAI-Voice-2.1: $22MAI-Voice-2.1-Flash: $15

The word error rate and latency figures are Microsoft’s own, published in the launch post, which claims the top spot on the Artificial Analysis leaderboard. At the time of writing the model does not yet appear on that board’s public streaming listing, where Grok Voice Transcribe 2.0 still sits first. Treat the ranking as unverified and the two speed figures as the reason to look.

Where you can run the MAI models

  • Microsoft Foundry, for production workloads inside Azure
  • MAI Playground, for trying it without wiring anything up
  • Vercel AI Gateway and OpenRouter, for teams already routing through either
  • LiveKit support announced as coming, with no date given
  • No open weights: all three models are hosted only

What Microsoft is not saying

Three gaps matter for anyone pricing a real deployment. There is no speaker attribution claim for the streaming model, so if you need to know who said what in a meeting, this does not answer it yet. There is no per language breakdown across the 60 languages, and a headline word error rate from clean English audio tells you nothing about Dutch on a mobile connection.

And there is no noisy or far field testing. Transcription scores collapse on room microphones and background noise, which is exactly where a call centre or a meeting recorder lives. The 0.23 point accuracy lead over Grok Voice Transcribe 2.0 works out at about two extra correct words per thousand. That is real, and it is not a reason to migrate anything.

What this means

Worth testing now if you are building a voice agent where the user waits on the reply. A four person team wiring a phone assistant on top of an LLM has two latency budgets: the model thinking and the audio pipeline around it. Cutting the transcript delay from roughly half a second to 0.13 seconds takes a noticeable pause out of every turn, and paired with MAI-Voice-2.1-Flash at 150 milliseconds you get a round trip that no longer feels like a walkie talkie. Test it on your own audio, in your own languages, not on the benchmark.

Worth watching, not switching, if you already run Deepgram or ElevenLabs and your transcripts are good enough. Microsoft is charging about 38% more per minute than either, and 5.4 times what its own batch model costs, for an accuracy difference you will struggle to notice. The one case where the sums clearly work is the opposite of the headline: if your audio is already finished when it reaches you, keep using the batch model at $0.10 an hour and spend nothing on this. Streaming is a product for conversations, and Microsoft is pricing it accordingly.

For more information, visit the official announcement of MAI-Transcribe-2-Streaming on the Microsoft AI blog.

Add DataNorth AI to your Google favorites