PulseAugur
中
实时 00:01:11
English(EN) Benchmarking Human and Automatic Speech Recognition of Diverse Speech: Initial Results

开源ASR模型多样化超越Whisper;比较人类与机器语音识别

一项新的研究论文比较了自动语音识别(ASR)系统和人工听众,发现Google Telephony表现良好,有时甚至在儿童或老年人语音以及弗拉芒语口音等特定语音类型上优于人类。与此同时,开源ASR模型领域正在超越Whisper实现多样化,Cohere的Transcribe和IBM的Granite Speech 4.1已成为强有力的竞争者。然而,这些模型的比较因评估方法不同和私有数据集的影响而变得复杂,这表明对于用户而言,许可、语言支持和成本正变得比原始准确性更关键的因素。 AI

影响 开源ASR模型的多元化以及与人类表现的比较,凸显了语音技术的进步和更广泛应用的潜力。

排序理由 该集群包含一项关于语音识别基准测试的研究论文以及对开源ASR模型的比较。

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 3 个来源。 我们如何撰写摘要 →

开源ASR模型多样化超越Whisper;比较人类与机器语音识别

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
该集群包含一项关于语音识别基准测试的研究论文以及对开源ASR模型的比较。
Source corroboration
3 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
model release, product, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
78 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+1 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准。

报道来源 [3]

  1. arXiv cs.CL TIER_1 English(EN) · Ilse Huisman, Rares Popa, Yuanyuan Zhang, Odette Scharenborg ·

    多样化语音的人工和自动语音识别基准测试:初步结果

    arXiv:2607.19049v1 Announce Type: new Abstract: Humans are often considered to be the best listeners and seen as the upper-bound performance of automatic speech recognition (ASR) systems. We present a preliminary comparison of the performances of state-of-the-art ASR systems and …

  2. MarkTechPost TIER_1 English(EN) · Asif Razzaq ·

    2026年最佳开源语音识别(ASR)模型:WER、语言、延迟和许可证对比

    <p>Open speech recognition stopped being a Whisper monoculture in 2026. Cohere Transcribe, IBM Granite Speech 4.1, ARK-ASR and MOSS-Transcribe are now separated by less than one WER point on the Hugging Face Open ASR Leaderboard — which means rank no longer decides anything. This…

  3. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    2026年,开放语音识别不再是Whisper一家独大。Cohere Transcribe、IBM Granite Speech 4.1、ARK-ASR和MOSS-Transcribe现已开创Hu

    Open speech recognition is no longer a Whisper monoculture in 2026. Cohere Transcribe, IBM Granite Speech 4.1, ARK-ASR and MOSS-Transcribe now trailblaze the Hugging Face Open ASR Leaderboard, separated by less than one WER point. A comparison of 16 open-weight models evaluates w…