PulseAugur
EN
LIVE 20:28:05

Open ASR Models Diversify Beyond Whisper; Human vs. Machine Speech Recognition Compared

A new research paper compares automatic speech recognition (ASR) systems with human listeners, finding that Google Telephony performed well, sometimes even outperforming humans on specific speech types like children's or older adults' speech and Flemish accents. Meanwhile, the open-source ASR model landscape is diversifying beyond Whisper, with Cohere's Transcribe and IBM's Granite Speech 4.1 emerging as strong contenders. However, the comparison of these models is complicated by differing evaluation methodologies and the impact of private datasets, suggesting that license, language support, and cost are becoming more critical factors than raw accuracy for users. AI

IMPACT The diversification of open ASR models and the comparison with human performance highlight advancements in speech technology and potential for wider adoption.

RANK_REASON The cluster includes a research paper on speech recognition benchmarking and a comparison of open-source ASR models.

Read on arXiv cs.CL →

AI-generated summary · Google Gemini · from 3 sources. How we write summaries →

Open ASR Models Diversify Beyond Whisper; Human vs. Machine Speech Recognition Compared

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
The cluster includes a research paper on speech recognition benchmarking and a comparison of open-source ASR models.
Source corroboration
3 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
model release, product, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
66 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+1 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

Full methodology in our editorial standards.

COVERAGE [3]

  1. arXiv cs.CL TIER_1 English(EN) · Ilse Huisman, Rares Popa, Yuanyuan Zhang, Odette Scharenborg ·

    Benchmarking Human and Automatic Speech Recognition of Diverse Speech: Initial Results

    arXiv:2607.19049v1 Announce Type: new Abstract: Humans are often considered to be the best listeners and seen as the upper-bound performance of automatic speech recognition (ASR) systems. We present a preliminary comparison of the performances of state-of-the-art ASR systems and …

  2. MarkTechPost TIER_1 English(EN) · Asif Razzaq ·

    Best Open Speech Recognition (ASR) Models in 2026: WER, Languages, Latency, and License Compared

    <p>Open speech recognition stopped being a Whisper monoculture in 2026. Cohere Transcribe, IBM Granite Speech 4.1, ARK-ASR and MOSS-Transcribe are now separated by less than one WER point on the Hugging Face Open ASR Leaderboard — which means rank no longer decides anything. This…

  3. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    Open speech recognition is no longer a Whisper monoculture in 2026. Cohere Transcribe, IBM Granite Speech 4.1, ARK-ASR and MOSS-Transcribe now trailblaze the Hu

    Open speech recognition is no longer a Whisper monoculture in 2026. Cohere Transcribe, IBM Granite Speech 4.1, ARK-ASR and MOSS-Transcribe now trailblaze the Hugging Face Open ASR Leaderboard, separated by less than one WER point. A comparison of 16 open-weight models evaluates w…