Microsoft launches MAI-Transcribe-2-Streaming

MAI-Transcribe-2-Streaming launches with transcripts in just over 100ms, alongside MAI-Voice-2.1 and Flash for multilingual voice agents.

· 2 min read
MAI

Microsoft AI has launched MAI-Transcribe-2-Streaming, a real-time speech-to-text model that the company says ranks first on Artificial Analysis for the accuracy of both final and partial transcripts. The release arrives alongside MAI-Voice-2.1 and MAI-Voice-2.1-Flash, forming an audio stack for developers building conversational voice agents.

The transcription model supports 60 languages with automatic, continuous language detection. It starts producing partial hypotheses just over 100 milliseconds after receiving audio, revises them as context arrives, and quickly commits stable text. This lets an agent begin reasoning or call tools before a speaker finishes, while live captions can appear during speech. Microsoft says its internal tests found words appearing twice as fast as with its closest competitor for real-time dictation and subtitling. The introductory price is $0.54 per audio hour through the end of the year.

MAI

MAI-Voice-2.1 brings Microsoft’s text-to-speech system to 23 languages and 26 locales. One speaker identity can switch among languages while retaining the same recognizable voice and adopting a native accent and local phrasing. This could let a tutoring app change languages without changing teachers or allow an assistant to answer in the language it hears. The model costs $22 per million characters.

MAI-Voice-2.1-Flash supports the same languages and cross-language speakers but targets high-volume, latency-sensitive workloads. Microsoft says it can generate 45 seconds of audio with end-to-end latency of 150 milliseconds, provides 55% faster model inference, and costs about 60% less than comparable models. Flash is priced at $15 per million characters. Both voice models can clone a voice across supported languages from a few seconds of reference audio and include consent guardrails intended to prevent misuse.

Microsoft AI is positioning the three-model lineup as a complete loop for agents that must hear, reason, use tools, and speak within a natural conversational window. Developers can access the two voice models through OpenRouter and all three through Microsoft Foundry, MAI Playground, Vercel, and Azure Voice Live, with LiveKit support coming soon. Microsoft has also released Chatter in MAI Playground to demonstrate the models working together.

  • MAI-Transcribe-2-Streaming developer overview: Related context: Microsoft's developer guide identifies the Foundry integration as public preview and explains access through an OpenAI Realtime-compatible WebSocket API or the Azure Speech SDK.

Source