AssemblyAI has released new benchmark data for its Universal-3.5 Pro and Universal-3.5 Pro Realtime speech-to-text models, reporting normalized word error rates (WER) of 4.35% and 5.53% respectively. The company emphasizes that WER alone is an insufficient metric for comparing vendors, highlighting the importance of metrics like concatenated minimum-permutation word error rate (cpWER) and Missed Entity Rate (MER) for conversational applications. AssemblyAI's data suggests a significant gap in entity recognition accuracy compared to other models on the Pipecat open STT benchmark, which evaluates streaming models on real voice agent conversations. AI
IMPACT Highlights the importance of specific metrics beyond word error rate for evaluating speech-to-text models in real-world applications.
RANK_REASON Blog post detailing performance metrics and benchmarks for existing speech-to-text models.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →