Microsoft has launched MAI-Transcribe-2-Streaming, a speech-recognition model that returns evolving text while someone is talking. Its October 1 announcement pairs the release with MAI-Voice-2.1 and the lower-latency MAI-Voice-2.1-Flash, giving developers separate components for hearing a request and speaking a reply.
In the official announcement, Microsoft says the transcription model supports 60 languages with automatic, continuous language detection. It reports a number-one accuracy position for both final and partial transcripts on Artificial Analysis. That ranking is Microsoft's account of the evaluation at launch, not a Franklin rerun or a guarantee for every language and recording condition.
Early partial transcripts can start work sooner
The model begins producing hypotheses, called partials, just over 100 milliseconds after receiving audio, according to Microsoft. It revises those hypotheses as more context arrives and then commits a stable transcript.
This sequence allows a live application to show words before a speaker finishes or begin preparing an agent response mid-sentence. It also introduces a distinction developers must handle: a partial transcript is still subject to change. Starting a reversible preparatory step is different from committing an external action on wording the recognizer may revise.
Microsoft reports that words appear twice as fast as with its closest competitor in internal dictation and subtitling evaluations. The captured announcement does not identify that competitor or fully describe those evaluation conditions. The comparison should therefore remain an attributed internal result, rather than an established speed ratio across products.
The introductory transcription price is $0.54 per hour of audio through the end of the year. That period applies to the launch offer; the article supplies no basis for assuming the rate continues unchanged afterward.
The voice models have different priorities
Microsoft describes MAI-Voice-2.1 as its strongest multilingual text-to-speech model to date, supporting 23 languages and 26 locales. The company says a single voice can speak across those languages while retaining its identity. That is a capability claim from the publisher, not independent confirmation of equal quality in every locale.
The standard voice model costs $22 per million characters. Flash costs $15 per million characters and is positioned for high-volume, latency-sensitive work. Microsoft reports that Flash can generate 45 seconds of audio with 150 milliseconds of end-to-end latency in the described measurement. Those durations refer to generated audio and generation delay respectively; they should not be conflated with the total delay of a conversational application.
The company also reports faster inference and lower cost than comparable models, but the announcement does not name the complete comparison set. A buyer needs specific alternatives and equivalent settings before using those percentages in a procurement calculation.
Both voice models accept a few seconds of reference audio for cloning across supported languages. Microsoft says they include consent guardrails. The announcement does not prove those safeguards prevent every misuse, and the availability of cloning does not create permission to reproduce another person's voice.
Availability and a demo make the parts accessible
Microsoft has built a MAI Playground demo called Chatter to show the models working in a live agent. The announcement lists Microsoft Foundry, MAI Playground, Vercel and Azure Voice Live as access routes for the models, with LiveKit marked as coming soon. It also lists the two voice models through OpenRouter.
The split between transcription and generation makes individual components available for evaluation. It leaves the application's reasoning, tool use and action policy to the developer. Lower delay at the audio endpoints can make more time available for those steps, but it does not establish that the complete agent will answer correctly or act safely.
An evaluation should measure partial-transcript revisions, recognition accuracy on representative recordings and the delay through a whole conversational turn. The launch provides concrete models and published pricing to begin that work, while its headline rankings and demonstrations remain evidence to inspect rather than outcomes to assume.