Microsoft releases streaming transcription and voice models
Microsoft released MAI-Transcribe-2-Streaming, ranked first on Artificial Analysis, plus two low-latency voice models.
Microsoft AI released MAI-Transcribe-2-Streaming on October 1, 2026, its first streaming speech-to-text model, alongside text-to-speech models MAI-Voice-2.1 and MAI-Voice-2.1-Flash. The transcription model covers 60 languages with continuous automatic language detection, returns initial hypotheses just over 100 milliseconds after audio arrives, and carries introductory pricing of $0.54 per audio hour through the end of 2026. MarkTechPost reports that Artificial Analysis ranks it first of 38 models, with 2.5% word error rate for first partials at 0.12 seconds and the same 2.5% rate 0.13 seconds after speech ends; The Decoder says it ranks first for accuracy and cites first partials in just over 100 milliseconds, without the model count or error-rate figures. MAI-Voice-2.1 speaks 23 languages in one voice, while MAI-Voice-2.1-Flash is cited at 150 milliseconds latency and $15 per million characters instead of $22. Both voice models can clone speech from a few seconds of reference audio, include misuse safeguards, and are available through Microsoft Foundry, MAI Playground, and OpenRouter. Weights are not open, and the public preview has no SLA.
- Microsoft AI released MAI-Transcribe-2-Streaming on October 1, 2026, described as its first streaming speech-to-text model, with MAI-Voice-2.1 and MAI-Voice-2.1-Flash.
- The transcription model covers 60 languages with continuous automatic language detection and emits first partials just over 100 ms after audio arrives.
- MarkTechPost says Artificial Analysis ranks it first of 38 models at 2.5% word error rate, including first partials at 0.12 seconds and final results 0.13 seconds after speech ends; The Decoder says it ranks first for accuracy without…
- Introductory transcription pricing is $0.54 per audio hour through the end of 2026; weights are not open and the public preview has no SLA.
- MAI-Voice-2.1 speaks 23 languages in one voice; MAI-Voice-2.1-Flash is cited at 150 ms latency and $15 per million characters instead of $22.
- Both voice models clone a voice from a few seconds of audio, include misuse safeguards, and are available through Microsoft Foundry, MAI Playground, and OpenRouter.
Coverage timelineoldest first · each row is one article
- · 6d agoMicrosoft AI releases new transcription and text-to-speech models for voice agents
The Decoder· 48
Microsoft released MAI-Transcribe-2-Streaming and two MAI-Voice text-to-speech models for low-latency voice agents.
- · 6d agoMicrosoft AI Releases MAI-Transcribe-2-Streaming: #1 Real-Time Speech-to-Text Model on Artificial Analysis
MarkTechPost· 72
Microsoft released MAI-Transcribe-2-Streaming, ranking first of 38 models on Artificial Analysis streaming speech-to-text accuracy.