PulseAugur
实时 09:52:08
English(EN) StreamHear: Domain-Adapted Pseudo-Labeling for Semi-Supervised Streaming Speech Recognition

StreamHear 管道增强了领域自适应流式语音识别

研究人员开发了 StreamHear,一个新颖的半监督学习管道,旨在提高流式自动语音识别 (ASR) 在领域偏移音频上的性能。该系统通过首先在可用的标记数据上微调一个离线变换器教师模型来适应预训练的流式学生模型。然后,该教师模型为未标记的音频生成伪标签,用于微调学生模型。此外,StreamHear 还包含一个重新对齐步骤,使用 ASR 假设锚点来优化词语位置,在各种数据集上均显示出持续的改进。 AI

影响 通过利用未标记数据提高专业音频的 ASR 性能,可能降低领域适应的成本。

排序理由 该集群包含一篇详细介绍语音识别新方法的论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

StreamHear 管道增强了领域自适应流式语音识别

报道来源 [1]

  1. arXiv cs.CL TIER_1 English(EN) · Zefang Liu, Chenyang Zhu, Sangwoo Cho, Xujun Peng, Shi-Xiong Zhang, Sambit Sahu ·

    StreamHear:领域自适应伪标签用于半监督流式语音识别

    arXiv:2608.13717v1 Announce Type: new Abstract: Streaming automatic speech recognition (ASR) underperforms on domain-shifted target audio, where labeled in-domain data is costly to prepare while unlabeled audio is abundant. We present StreamHear, a semi-supervised pipeline that a…