AI·News & analysis
Microsoft launches real-time transcription and new voice models for AI voice agents
MAI-Transcribe-2-Streaming turns speech into text as people talk, and MAI-Voice-2.1 and its Flash version talk back. Together, Microsoft pitches them as the building blocks for voice agents with fewer awkward pauses.

Tide
Ripple
Sci-fi
5/10
Reality
In your hands
Captions that appear before you finish your sentence.How we rate
Microsoft AI launched three audio models for voice agents: MAI-Transcribe-2-Streaming, which turns speech into text in real time across 60 languages, and two text-to-speech models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash.
Microsoft says the transcription model tops a public accuracy ranking. All three are available now through Microsoft Foundry and partners, priced per audio hour or per million characters.
What to know
- Microsoft AI launched MAI-Transcribe-2-Streaming, its first streaming transcription model, plus two text-to-speech models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash.
- The transcription model works in 60 languages and sends early text in just over 100 milliseconds, Microsoft says; it costs $0.54 per audio hour.
- MAI-Voice-2.1 covers 23 languages and 26 locales at $22 per million characters; the Flash version costs $15 and targets about 150 milliseconds of latency.
- All three are available through Microsoft Foundry, the MAI Playground, Vercel and Azure Voice Live; Microsoft's performance figures are its own.
Microsoft launched three new audio models built for AI voice agents: one that listens in real time and two that talk back. The interesting question: why does a few hundred milliseconds matter so much in a conversation with a computer?
What did Microsoft launch?
Microsoft AI released three models on October 1, according to its announcement:
- MAI-Transcribe-2-Streaming: its first streaming transcription model, which turns speech into text while you're still talking.
- MAI-Voice-2.1: a text-to-speech model for expressive, high-quality voices.
- MAI-Voice-2.1-Flash: a faster, cheaper version of the voice model.
Microsoft pitches them as "building blocks" for conversational voice agents, the kind of AI that answers phone calls, tutors students or runs a voice assistant. All three were available the same day, RuntimeWire reports.
Why does real-time transcription matter?
Most transcription waits until you finish a sentence, then returns the text. That creates a pause before a voice agent can even start thinking.
The streaming model works differently. RuntimeWire explains that it takes in continuous audio and sends incremental transcript updates before confirming a final version. Unite.AI reports Microsoft's numbers: the model produces its first guesses, called partials, in just over 100 milliseconds after receiving audio, then revises them as more words arrive.
Microsoft says this lets voice agents start reasoning or calling tools mid-sentence, before the speaker finishes.
Microsoft frames the pairing of its streaming transcription and Flash voice models as buying back time on both ends of the voice-agent loop: hearing, understanding, deciding and speaking, all within the short window where a person still feels like they're having a conversation, Unite.AI reports.
Think of it like a good friend who starts nodding before you finish your sentence. They don't wait for the period to start understanding you.
For live dictation and subtitles, Microsoft says its internal tests show words appearing twice as fast as with its closest competitor, according to Unite.AI.
How accurate is it?
Microsoft says the model ranks No. 1 for accuracy on the Artificial Analysis streaming leaderboard, for both final and partial transcripts.
Here are the leaderboard figures Microsoft cited, as reported by Unite.AI, from a September 28, 2026 snapshot:
- Final word-error rate: 2.5%
- First-partial word-error rate: 2.8%
- Time to final transcript: 0.13 seconds
Word-error rate is the share of words a model gets wrong, so lower is better. Unite.AI says the Artificial Analysis test uses about eight hours of real-world audio with varied accents, specialist vocabulary and tough sound conditions.
The model supports 60 languages with automatic, continuous language detection, so it can follow a speaker who switches languages.
Microsoft also says the model sits on the benchmark's Pareto frontier for accuracy versus speed. In plain terms, you don't have to give up much speed to get its accuracy. Coverage of the launch cited by RuntimeWire reported words appearing as early as 320 milliseconds in most cases.
It builds on MAI-Transcribe-2, Microsoft's earlier non-streaming speech model, which the company billed as the fastest, most accurate and cheapest in the world, Unite.AI notes.
The catch: RuntimeWire points out that these are Microsoft's reported model figures. They don't guarantee how fast a real voice agent responds, which also depends on networks, the app and the system writing the reply. And the announcement alone doesn't show the models beat rivals in independent, like-for-like tests.
What about the voice models?
The two MAI-Voice models handle the other half: turning text into speech.
MAI-Voice-2.1 supports 23 languages and 26 locales, according to Unite.AI. Microsoft says a single voice can speak all of them with a native accent and keep the same identity when switching. Its example: a tutoring app that changes languages mid-lesson without swapping teachers.
MAI-Voice-2.1-Flash covers the same languages but is built for high-volume, speed-sensitive work. Microsoft says it can generate up to 45 seconds of audio with about 150 milliseconds of end-to-end latency. It also claims Flash is about 55% faster at inference and roughly 60% cheaper than comparable models.
Two more claims from Microsoft, via Unite.AI:
- Voice cloning: both models can clone a voice from a few seconds of reference audio, with built-in consent guardrails that Microsoft says prevent misuse.
- Human-like: in a 4,000-listener test, 50.3% of listeners rated the MAI voices as equally or more human-like than real human recordings.
What does it cost?
Here's the pricing reported by RuntimeWire and Unite.AI:
| Model | Price |
|---|---|
| MAI-Transcribe-2-Streaming | $0.54 per audio hour (introductory, through end of year) |
| MAI-Voice-2.1 | $22 per million characters |
| MAI-Voice-2.1-Flash | $15 per million characters |
RuntimeWire puts the transcription price in context. Microsoft's batch model, MAI-Transcribe-2, launched in September at $0.10 per audio hour, so streaming costs 5.4 times more. That's not a perfect comparison, since streaming has to keep sending provisional results while audio is still arriving. RuntimeWire also notes the $0.10 batch price was a limited-time offer through the end of 2026.
Where can you use them?
Microsoft lists several places, according to Unite.AI:
- All three models: Microsoft Foundry, the MAI Playground, Vercel and Azure Voice Live.
- Voice models: also on OpenRouter.
- Coming soon: LiveKit.
Microsoft also built a demo called Chatter in the MAI Playground, where you can talk to a voice assistant powered by the new models.
Microsoft's suggested uses include customer service agents that transcribe requests as they're spoken and answer in natural speech, multilingual assistants that reply in whatever language they're addressed in, and learning apps with distinct voices for tutoring, role-play and narration.
In real life You call a store's support line and say your order is late. A voice agent built on these models starts working on your request while you're still talking, then answers in a natural voice, in your language, without that long robotic pause.
Who's behind it?
The models come from Microsoft AI, led by CEO Mustafa Suleyman. RuntimeWire notes Suleyman co-founded DeepMind before Google bought it, then co-founded Inflection AI, which built the Pi assistant. He joined Microsoft in 2024 to lead Microsoft AI, a group formed with Inflection co-founder Karén Simonyan and several Inflection employees.
RuntimeWire also explains why selling the pieces separately matters. A voice agent has to recognize speech, decide what to do and speak a reply. Separate models for listening and speaking let developers pick their own trade-offs between quality, speed and cost for each step.
The big picture: RuntimeWire sees a clear commercial bet. As more apps add voice, Microsoft wants its own models in the conversation, not just its cloud as the place apps run. It's also been busy elsewhere, such as dropping the Copilot+ PC brand.
What it means for you
- If you build voice apps: you can try all three models today in Microsoft Foundry or the MAI Playground, and compare streaming against batch pricing.
- If you use voice assistants or call centers: future versions may feel faster and more natural, with fewer awkward pauses.
- If you speak several languages: one assistant voice could switch languages mid-conversation without sounding like a different person.
- Watch the fine print: speed and accuracy figures are Microsoft's own, and real apps may perform differently.
The bottom line
Microsoft released a real-time transcription model and two voice models designed to make AI voice agents respond faster and sound more natural. The prices are clear and the models are available now. Whether they beat rivals in real use is something independent tests still need to show.
Key facts
- Models
- MAI-Transcribe-2-Streaming, MAI-Voice-2.1, MAI-Voice-2.1-Flash
- Transcription languages
- 60, with automatic language detection
- Transcription price
- $0.54 per audio hour (introductory)
- Voice prices
- $22 (MAI-Voice-2.1) and $15 (Flash) per million characters
- Voice languages
- 23 languages, 26 locales
Got questions?
Quick answers, plain wordsWhat is MAI-Transcribe-2-Streaming?
It's Microsoft's first streaming transcription model. It turns speech into text while someone is still talking, sending early guesses and then a final transcript, in 60 languages.
What are MAI-Voice-2.1 and MAI-Voice-2.1-Flash?
They're Microsoft's new text-to-speech models. MAI-Voice-2.1 aims for expressive, high-quality speech, while Flash is built for speed and lower cost.
How much do Microsoft's MAI voice models cost?
The streaming transcription model costs $0.54 per hour of audio. MAI-Voice-2.1 costs $22 per million characters and Flash costs $15.
How fast is MAI-Transcribe-2-Streaming?
Microsoft says it produces its first partial transcript in just over 100 milliseconds of receiving audio. Real-world speed also depends on networks and the app.
Can MAI-Voice clone voices?
Yes. Microsoft says both voice models can clone a voice from a few seconds of reference audio, with built-in consent guardrails meant to prevent misuse.
Where can developers use these models?
All three are available through Microsoft Foundry, the MAI Playground, Vercel and Azure Voice Live. The voice models are also on OpenRouter, and LiveKit support is listed as coming soon.
How many languages does MAI-Voice-2.1 support?
23 languages and 26 locales. Microsoft says one voice can speak all of them with a native accent while keeping the same speaker identity.
Who leads Microsoft AI?
Mustafa Suleyman, who co-founded DeepMind and Inflection AI before joining Microsoft in 2024 to lead Microsoft AI, according to RuntimeWire.
Is MAI-Transcribe-2-Streaming really the most accurate?
Microsoft says it ranks No. 1 for accuracy on the Artificial Analysis streaming leaderboard. RuntimeWire notes the launch itself doesn't prove it beats rivals in independent, like-for-like tests.
SourcesMicrosoft AI
Topics and tagsMicrosoft, AI agents, microsoft, speech to text
Related stories

Bill Gates warns AI is 'powerful enough' to cause a billion deaths
The Microsoft co-founder says AI could drive catastrophic harm if misused, and called for mandatory government oversight rather than relying on tech companies to self-regulate.

Cloudflare releases Clef, open-weight AI models that make yes-or-no decisions fast
Clef and Clef-flash answer bounded questions for AI agents in milliseconds instead of writing text. Cloudflare hosts them and also gives away the weights, a direct challenge to TypeSafe's Jev.

ChatGPT can now show you wearing clothes before you buy them
OpenAI launched virtual try-on and a Favorites list in ChatGPT's shopping results worldwide. Upload a photo of yourself and see how a jacket or accessory might look on you.
More in brief
- California will fine robotaxi companies that block first responders for over 30 minutesOct 2
- Apple's smart home hub reportedly launches October 13, with a camera that never records videoOct 1
- GrayKey maker reportedly found a way around the iPhone's Inactivity RebootOct 1
- Fervo's Cape Station becomes the first enhanced geothermal plant to sell power commerciallyOct 1
- Cloudflare renames its data platform Basin and makes it generally availableOct 1
- Judge dismisses Chegg and Penske antitrust suits over Google's AI OverviewsOct 1