Wain AI/Tech Blog

AI news and trends worldwide, updated nearly every day

Microsoft Releases MAI-Transcribe-2 Speech Recognition - $0.10 per Hour of Audio, Priced Through Year-End

Image: Microsoft AI official website

Microsoft Releases MAI-Transcribe-2 Speech Recognition - $0.10 per Hour of Audio, Priced Through Year-End

Microsoft AI released MAI-Transcribe-2 on September 3, 2026. It reports a 5.2% average word error rate across 60 FLEURS languages, first in the company's own comparison, at $0.10 per hour of audio. That price is a limited-time offer through the end of the year, and on the two Artificial Analysis charts on the same page the model places second.

Microsoft AI released a speech recognition model called MAI-Transcribe-2 on September 3, 20261. The company positions it as “not only our most capable transcription model yet, but the most capable and efficient amongst our competitors”1.

The number that stands out is the price: $0.10 per hour of audio1. This is not a permanent change — the post says that at launch the model “will be priced at $0.10 per hour as a limited-time offer until the end of the year”1.

5.2% across 60 languages, 1.8% for Japanese

Microsoft cites three results: first place on the FLEURS benchmark across 60 languages with an average word error rate (WER) of 5.2%, defining the Pareto frontier for accuracy and latency on Artificial Analysis, and second place on the Artificial Analysis word error rate leaderboard1.

The FLEURS comparison is given two ways: with the language specified (forced language) and with the model inferring it (inferred language)1.

ModelForced languageInferred language
MAI-Transcribe-25.2%5.2%
Gemini 3.1 Pro5.3%5.8%
Gemini 3.5 Transcribe5.9%8.6%
Scribe v26.2%6.5%
Gemini 3.6 Flash6.9%9.4%
GPT-transcribe10.4%10.6%
Whisper v3-large22.8%23.5%

The right-hand column is where the gap opens. MAI-Transcribe-2 stays at 5.2% whether or not the language is given, while Gemini 3.5 Transcribe, which entered public preview on August 26, goes from 5.9% to 8.6%, and Gemini 3.6 Flash from 6.9% to 9.4%1. For audio where the language is not known in advance, that difference matters. (Footnotes on the chart state that Urdu for Gemini 3.5 Transcribe, and Bulgarian and Swahili for Gemini 3.6 Flash, use Gemini 3.1 Pro numbers1.)

A per-language breakdown is also published, measured with the language specified1. Japanese comes in at 1.8%, the lowest of the seven models listed1. Chinese is 4.5%, the lowest, and Italian is 1.1%, tied with Gemini 3.5 Transcribe for lowest1.

It does not win everywhere. English is 2.7%, behind the 2.6% of both Gemini 3.1 Pro and Scribe v21. Urdu is 13.3%, above Gemini 3.1 Pro’s 12.1%1. Hebrew is 13.7% — the lowest of the seven, but all seven are in double digits1. Whether your workload is mostly English or spans other languages changes the conclusion.

Three superlatives in the headline, and the charts on the same page

The post’s title calls this “the fastest, most accurate and cheapest speech recognition model in the world”1. The charts on that same page do not go that far.

On the Artificial Analysis word error rate index (AA-WER) chart the company publishes, MAI-Transcribe-2 sits at 2.0%, in second place1. First is Alibaba’s Fun-Realtime-ASR-preview at 1.7%1. Behind them come ElevenLabs’ Scribe v2 at 2.2%, Microsoft’s own previous generation MAI-Transcribe-1.5 at 2.4%, Gemini 3.5 Transcribe at 2.6%, OpenAI’s GPT Transcribe at 3.3% and Whisper Large v3 at 4.1%1. The index is noted as combining three datasets: AA-AgentTalk (50%), VoxPopuli-Cleaned-AA (25%) and Earnings22-Cleaned-AA (25%)1. The body text itself says the model “ranks second” there1.

On the speed factor chart (seconds of input audio transcribed per second), first place goes to Deepgram’s Nova-3 at 501.4, with MAI-Transcribe-2 second at 403.61. All of these figures come from the charts on Microsoft’s page; the Artificial Analysis leaderboard updates continuously, so the values shown there now may differ. Behind them are Reson8’s Resonant-1 at 331.2, Whisper Large v3 via together.ai at 287.9, SpaceXAI’s Grok Speech to Text at 225, Gemini 3.5 Transcribe at 79.6, Scribe v2 at 54.7 and GPT Transcribe at 40.81. Microsoft notes that at a speed factor of 400, an hour of audio comes back in about nine seconds1.

The multiples Microsoft quotes in the text, attributed to evaluations run by Artificial Analysis, are “10x faster than OpenAI’s GPT-Transcribe, 7x faster than ElevenLabs’ Scribe v2, and 5x faster than Gemini 3.5 Transcribe while delivering higher accuracy”1. Against the speed factor figures above, those work out to roughly 9.9x, 7.4x and 5.1x, which is consistent.

So what the page’s own material shows is first place on the FLEURS 60-language average, plus a second place on each of accuracy and speed. Microsoft describes its scatter plot as placing MAI-Transcribe-2 alone in the most attractive quadrant1, but no single metric here shows it first in the world. As the tests published by Hume AI researchers in August showed, speech recognition benchmarks can be reproduced right down to the errors in the reference transcripts. Numbers a vendor lines up are worth reading down to which metric under which conditions.

What the model actually does

On features, most of what is listed sits around the transcription rather than in it1.

  • Speaker diarization — distinguishes speakers and attributes words to the right person within a recording
  • Word-level timestamps — precise timing for every word, for alignment, search, navigation and editing
  • Keyword biasing — helps the model recognize domain-specific terminology, abbreviations and names that are hard to pick out from context alone
  • Configurable transcription styles — “verbatim” captures speech as spoken including filler words and false starts; “clean” removes fillers for more readable captions and notes
  • Code switching — handles conversations that move between languages, including blends such as Hinglish and Spanglish
  • Automatic language identification — detects the language being spoken without the user specifying it in advance

The applications Microsoft names are clinical note-taking, legal documentation, accessibility and closed captioning1. Seen against that list, the verbatim/clean split makes sense: compliance and analysis workloads need the false starts kept, while published captions do not.

The model is available to demo from launch day through Microsoft Foundry, MAI Playground and OpenRouter1.

What an order-of-magnitude price change moves

In work that processes audio in volume — meeting minutes, call center records, video captioning — the per-hour cost of transcription is often what caps how much gets processed at all. A rate of $0.10 per hour leaves room to revisit setups where only part of the material was being run.

Two conditions attach. One is that this price is explicitly a limited-time offer through the end of the year1. Treating it as a permanent assumption in a budget that crosses into next year, or in a multi-year estimate, is risky. The other is the spread across languages. The 5.2% average over 60 languages flattens a range that runs from 1.8% for Japanese to 13.7% for Hebrew1. For a workload concentrated in particular languages, the per-language figures describe reality better than the average.

Microsoft has been building out the MAI line with purpose-specific models — MAI-Code-1.1-Flash for coding and MAI-Cyber-1-Flash for security among them. Speech recognition follows the same pattern: a dedicated model offered cheaply and quickly, rather than routing audio through a large general-purpose model. On the evaluation side, material for measuring per-language performance is growing too, as with the addition of Hindi and Indian English test sets to the Open ASR Leaderboard. Comparing a vendor’s average against a measurement on your own audio is getting easier to do.

Sources

  1. MAI-Transcribe-2 is the fastest, most accurate and cheapest speech recognition model in the world - Microsoft AI blog (September 3, 2026). Body text, feature list and the figures in the published charts (FLEURS 60-language average, AA-WER index, speed factor, per-language WER)

We publish the latest AI news nearly every day.

Subscribe via RSS Get new posts the moment they go live.

Search other keywords →