English(EN)How to Leverage Synthetic Speech for LLM-Based ASR Systems?
新研究推动了用于构音障碍语音和合成数据使用的ASR · 跟踪4个来源
作者PulseAugur 编辑部·[6 个来源]·
研究人员正在探索改进自动语音识别(ASR)系统的新方法。一项研究详细介绍了如何使用个性化数据对Whisper模型进行微调,显著降低了构音障碍语音的词错误率,在使用大量数据的情况下达到了9.7%的错误率。另一篇论文研究了使用合成语音训练ASR系统,发现通过房间冲激响应增强合成音频可以弥合与真实世界数据的差距。此外,一个名为PreferenceASR的新测试集已被开发出来,用于根据ASR系统遵循用户指定输出偏好的能力进行评估,揭示了传统基准测试所掩盖的性能差异。
AI
arXiv:2607.01733v1 Announce Type: new Abstract: Speech-LLM integration has shown promising results by leveraging extensive textual pretraining, yet its specific benefits for automatic speech recognition (ASR) remain unclear. We observe that as supervised ASR training data increas…
Speech-LLM integration has shown promising results by leveraging extensive textual pretraining, yet its specific benefits for automatic speech recognition (ASR) remain unclear. We observe that as supervised ASR training data increases, the contribution of LLM priors becomes less …
arXiv cs.CL
TIER_1English(EN)·Christian Huber, Laura Kernahan, Alexander Waibel·
arXiv:2606.31722v1 Announce Type: new Abstract: Automatic speech recognition (ASR) systems often perform poorly in dysarthric speech, limiting their usefulness to affected speakers in everyday communication. This paper presents a personalized ASR system for a dysarthric speaker, …
Automatic speech recognition (ASR) systems often perform poorly in dysarthric speech, limiting their usefulness to affected speakers in everyday communication. This paper presents a personalized ASR system for a dysarthric speaker, built by adapting a foundation ASR model to spea…
arXiv cs.AI
TIER_1English(EN)·Yanis Labrak, Dairazalia Sanchez-Cortes, Sergio Burdisso, S\'everin Baroudi, Shashi Kumar, Esa\'u Villatoro-Tello, Srikanth Madikeri, Manjunath K E, Old\v{r}ich Plchot, Kadri Hacio\u{g}lu, Petr Motlicek, Andreas Stolcke·
arXiv:2606.29031v1 Announce Type: cross Abstract: In regulated domains such as banking and healthcare, where privacy constraints make real speech costly to collect and retain, synthetic speech from modern text-to-speech (TTS) is an appealing alternative for training automatic spe…
arXiv cs.CL
TIER_1English(EN)·Nithin Rao Koluguri, Sasha Meister, Nikolay Karpov, Piotr Zelasko, Desh Raj, Jagadeesh Balam, Boris Ginsburg·
arXiv:2606.29534v1 Announce Type: new Abstract: Popular ASR test sets adopt inconsistent conventions for numbers, disfluencies, entities, and casing, while standard normalizers erase the format distinctions users care about. Current benchmarks therefore cannot measure whether a m…