Microsoft AI has launched MAI-Transcribe-2-Streaming, its first real-time speech-to-text model, on October 1, 2026. The model ranks first among 38 competitors on Artificial Analysis for both final and initial partial transcript accuracy, achieving a 2.5% Word Error Rate (WER) with minimal latency. It supports 60 languages and offers continuous language detection, enabling agents to process information mid-sentence. The introductory pricing is set at $0.54 per hour, with integration available via API or the MAI Playground. AI
IMPACT This model's high accuracy and low latency could significantly improve voice agents and real-time transcription services.
RANK_REASON Microsoft AI lab released a new speech-to-text model with benchmark performance claims.
Read on Mastodon — mastodon.social →
- Artificial Analysis
- Azure Speech SDK
- Cartesia Ink-2
- Grok Voice Transcribe 2.0
- MAI Playground
- MAI-Transcribe-2-Streaming
- MAI-Voice-2.1
- MAI-Voice-2.1-Flash
- Meta
- Microsoft AI
- Muse Voice Transcribe
- OpenAI
AI-generated summary · Google Gemini · from 4 sources. How we write summaries →