PulseAugur
实时 22:32:43
English(EN) Robust Assamese Speech Recognition through Controlled Fine-Tuning of Whisper Models

Whisper模型微调以实现鲁棒的阿萨姆语语音识别

研究人员开发了一个微调版本的Whisper模型,以改进阿萨姆语的自动语音识别(ASR)。该微调模型在Mozilla Common Voice 24.0-Assamese语料库上进行训练,并使用Tesla 4 GPU针对资源受限环境进行了优化,其性能显著优于零样本基线。新系统在词错误率(WER)、字符错误率(CER)、匹配错误率(MER)和词信息丢失(WIL)方面实现了大幅降低,同时在语义评估的BLEU和METEOR分数方面也有显著提高。 AI

影响 提高低资源语言语音技术的可访问性,可能催生新的应用。

排序理由 该集群包含一篇学术论文,详细介绍了使用微调的现有模型改进低资源语言语音识别的新方法。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.LG 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Whisper模型微调以实现鲁棒的阿萨姆语语音识别

报道来源 [1]

  1. arXiv cs.LG TIER_1 English(EN) · Ganapati Das, Dwipen Laskar, Hasin Afzal Ahmed, Sanjib Kr Kalita, Kshirod Sarmah, Hem Chandra Das, Manjula Kalita ·

    通过对Whisper模型进行受控微调实现鲁棒的阿萨姆语语音识别

    arXiv:2607.17164v1 Announce Type: new Abstract: Developing Automatic Speech Recognition (ASR) for morphologically rich, low-resource languages such as Assamese is challenging due to insufficient annotated speech data. The pretrained Whisper model performs poorly on Assamese speec…