Microsoft Released MAI-Transcribe-2

04-09-2026

MAI-Transcribe-2 pairs a 2.0% independent AA-WER score with roughly 410x real-time processing and a promotional $0.10 per audio hour. Built-in diarization, timestamps and keyword biasing make it a serious option for high-volume transcription.

Written by:

Jorick van Weelie

Marketing Lead at DataNorth | AI Enthusiast & Tech Storyteller

microsoft mai transcribe 2 price, speed and accuracy
Sign up for our Newsletter

Published 4 September 2026

Microsoft released MAI-Transcribe-2 on 3 September 2026 at a promotional price of $0.10 per audio hour. The model adds speaker diarization, word-level timestamps, domain keyword biasing and 60-language support. More importantly, independent Artificial Analysis measurements support the core commercial pitch: MAI-Transcribe-2 combines 2.0% word error rate with roughly 410.7 times real-time processing and an estimated $1.67 per 1,000 minutes.

For call-center analytics, meeting intelligence, captioning and clinical or legal transcription, that combination is worth attention now. Microsoft has improved accuracy and feature coverage over MAI-Transcribe-1.5 while cutting the launch price by more than 70%. The main uncertainty is economic rather than technical: $0.10 per hour is explicitly a limited-time offer through the end of 2026, and Microsoft has not published the permanent price.

What MAI-Transcribe-2 changes for transcription teams

MAI-Transcribe-2 is aimed at production workflows that normally require several layers around a speech model. Speaker diarization identifies who said what, word-level timestamps support search and editing, and keyword biasing helps with product names, medical terms, abbreviations and other domain language. Developers can also choose a verbatim style that preserves fillers and false starts or a clean style that removes them for readable notes and captions.

The model automatically identifies languages and supports code switching, including mixed-language conversations such as Hinglish and Spanglish. Microsoft says it is robust to noisy audio and covers 60 languages, up from 43 in MAI-Transcribe-1.5. That matters for teams currently routing calls to different language-specific systems or adding post-processing for speaker labels and timestamps.

Microsoft makes the model available through Microsoft Foundry using Azure Speech, the MAI Playground and OpenRouter. The model page says one hour of audio requires about 10 seconds of model inference, compared with 20 seconds for MAI-Transcribe-1.5.

Accuracy, speed and price compared with rivals

Microsoft’s launch headline calls MAI-Transcribe-2 the fastest, most accurate and cheapest speech recognition model in the world. The independent data supports a strong combined position, but not literal leadership on every individual metric. Artificial Analysis currently measures MAI-Transcribe-2 at 2.0% AA-WER and a 410.7x speed factor. Deepgram Nova-3 is faster on raw throughput at around 600x, but its AA-WER is materially higher. ElevenLabs Scribe v2 is close on accuracy at 2.2% but much slower.

MeasureMAI-Transcribe-2Relevant comparison
Artificial Analysis AA-WER2.0%Scribe v2 2.2%; MAI-Transcribe-1.5 2.4%
Artificial Analysis speed factor410.7xNova-3 about 600x; Scribe v2 about 53x
Estimated cost per 1,000 minutes$1.67Scribe v2 $3.67; Nova-3 $4.30
FLEURS, top 25 languages3.4% WERMAI-Transcribe-1.5 3.7%
Supported languages60MAI-Transcribe-1.5 43
1 hour audio, model inferenceabout 10 secMAI-Transcribe-1.5 about 20 sec

Artificial Analysis calculates AA-WER across roughly eight hours of speech from three datasets: AgentTalk, VoxPopuli-Cleaned-AA and Earnings22-Cleaned-AA. Its speed factor is the median amount of audio transcribed per second of processing time, based on 10-minute files. That makes the figures useful for relative comparison, but they do not replace testing with your own accents, channels, vocabulary and recording conditions.

What Microsoft still has to prove

The biggest missing number is the price after 31 December 2026. At the promotional $0.10 per hour, the business case is unusually strong. If the permanent price rises materially, the comparison with Scribe v2, Gemini Transcribe, Deepgram and specialist providers changes. Procurement teams should therefore treat the launch rate as a pilot price, not a long-term TCO assumption.

The benchmark story also needs care. Microsoft reports a 5.2% average WER across all 60 FLEURS languages, while its model page highlights 3.4% for the top 25 languages. Aggregate multilingual scores can hide weak languages, accents or domains. A European contact center should test its actual language mix, cross-talk and telephony compression before replacing an existing pipeline.

Finally, raw speed is not the same as end-to-end latency. Artificial Analysis shows Nova-3 with higher throughput than MAI-Transcribe-2, despite Microsoft’s broad ‘fastest’ wording. Microsoft’s stronger claim is the combination of low error rate, high batch speed and low launch price, not first place on every single axis.

What this means

For teams processing large volumes of recorded speech, MAI-Transcribe-2 is worth testing now. The first pilot should use difficult production audio: noisy calls, multiple speakers, specialist vocabulary and the languages that matter commercially. Measure WER, speaker attribution, timestamp accuracy, correction time and effective cost per usable transcript.

The release is particularly compelling for organizations that currently pay separately for transcription and post-processing features. The built-in diarization, timestamps and contextual biasing can simplify that stack. The verdict is worth testing now, with one procurement caveat: do not lock a 2027 business case to the $0.10 promotional rate until Microsoft publishes permanent pricing.

For more information, visit the official announcement of MAI-Transcribe-2 on the Microsoft website.

Add DataNorth AI to your Google favorites