PulseAugur
实时 03:51:06
English(EN) How to Leverage Synthetic Speech for LLM-Based ASR Systems?

新研究推动了用于构音障碍语音和合成数据使用的ASR · 跟踪4个来源

研究人员正在探索改进自动语音识别(ASR)系统的新方法。一项研究详细介绍了如何使用个性化数据对Whisper模型进行微调,显著降低了构音障碍语音的词错误率,在使用大量数据的情况下达到了9.7%的错误率。另一篇论文研究了使用合成语音训练ASR系统,发现通过房间冲激响应增强合成音频可以弥合与真实世界数据的差距。此外,一个名为PreferenceASR的新测试集已被开发出来,用于根据ASR系统遵循用户指定输出偏好的能力进行评估,揭示了传统基准测试所掩盖的性能差异。 AI

影响 ASR个性化和合成数据利用方面的进步可以拓宽语音技术对不同用户群体的可及性。

排序理由 该集群包含多篇在arXiv上发表的学术论文,详细介绍了对ASR系统的研究。

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 6 个来源。 我们如何撰写摘要 →

新研究推动了用于构音障碍语音和合成数据使用的ASR · 跟踪4个来源

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
该集群包含多篇在arXiv上发表的学术论文,详细介绍了对ASR系统的研究。
Source corroboration
6 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
paper, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
71 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+2 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准

报道来源 [6]

  1. arXiv cs.CL TIER_1 English(EN) · Ruchao Fan, Yiming Wang, Rui Zhao, Liliang Ren, Keqi Deng, Xiaoyang Chen, Ali Zare, Bo Ren, Yuxuan Hu, Junkun Chen, Yan Huang, Yelong Shen, Jinyu Li ·

    重新思考语音-LLM集成在ASR中的应用:通过交错实现有效的联合语音-文本训练

    arXiv:2607.01733v1 Announce Type: new Abstract: Speech-LLM integration has shown promising results by leveraging extensive textual pretraining, yet its specific benefits for automatic speech recognition (ASR) remain unclear. We observe that as supervised ASR training data increas…

  2. arXiv cs.CL TIER_1 English(EN) · Jinyu Li ·

    重新思考语音-LLM集成在ASR中的应用:通过交错实现有效的联合语音-文本训练

    Speech-LLM integration has shown promising results by leveraging extensive textual pretraining, yet its specific benefits for automatic speech recognition (ASR) remain unclear. We observe that as supervised ASR training data increases, the contribution of LLM priors becomes less …

  3. arXiv cs.CL TIER_1 English(EN) · Christian Huber, Laura Kernahan, Alexander Waibel ·

    将基础语音识别模型应用于构音障碍语音:一个案例研究

    arXiv:2606.31722v1 Announce Type: new Abstract: Automatic speech recognition (ASR) systems often perform poorly in dysarthric speech, limiting their usefulness to affected speakers in everyday communication. This paper presents a personalized ASR system for a dysarthric speaker, …

  4. arXiv cs.CL TIER_1 English(EN) · Alexander Waibel ·

    将基础ASR模型适配到构音障碍语音:一个案例研究

    Automatic speech recognition (ASR) systems often perform poorly in dysarthric speech, limiting their usefulness to affected speakers in everyday communication. This paper presents a personalized ASR system for a dysarthric speaker, built by adapting a foundation ASR model to spea…

  5. arXiv cs.AI TIER_1 English(EN) · Yanis Labrak, Dairazalia Sanchez-Cortes, Sergio Burdisso, S\'everin Baroudi, Shashi Kumar, Esa\'u Villatoro-Tello, Srikanth Madikeri, Manjunath K E, Old\v{r}ich Plchot, Kadri Hacio\u{g}lu, Petr Motlicek, Andreas Stolcke ·

    如何利用合成语音为基于LLM的ASR系统赋能?

    arXiv:2606.29031v1 Announce Type: cross Abstract: In regulated domains such as banking and healthcare, where privacy constraints make real speech costly to collect and retain, synthetic speech from modern text-to-speech (TTS) is an appealing alternative for training automatic spe…

  6. arXiv cs.CL TIER_1 English(EN) · Nithin Rao Koluguri, Sasha Meister, Nikolay Karpov, Piotr Zelasko, Desh Raj, Jagadeesh Balam, Boris Ginsburg ·

    Preference-ASR:在语音大模型时代用于基准测试ASR的偏好感知测试集

    arXiv:2606.29534v1 Announce Type: new Abstract: Popular ASR test sets adopt inconsistent conventions for numbers, disfluencies, entities, and casing, while standard normalizers erase the format distinctions users care about. Current benchmarks therefore cannot measure whether a m…