Microsoft AI Releases MAI-Transcribe-2-Streaming: #1 Real-Time Speech-to-Text Model on Artificial Analysis
Microsoft released MAI-Transcribe-2-Streaming, ranking first of 38 models on Artificial Analysis streaming speech-to-text accuracy.
Microsoft AI released MAI-Transcribe-2-Streaming on October 1, 2026, its first streaming speech-to-text model, alongside TTS models MAI-Voice-2.1 and MAI-Voice-2.1-Flash. Artificial Analysis ranks it first of 38 models, with 2.5% word error rate 0.13 seconds after speech ends and the same 2.5% WER for first partials at 0.12 seconds. It covers 60 languages with continuous automatic language detection and emits initial hypotheses just over 100ms after audio arrives. Introductory pricing is $0.54 per audio hour through the end of 2026; weights are not open and the public preview has no SLA.