PulseAugur
实时 09:57:28

Speech-LLM 系统在 MLC-SLM 挑战赛中取得高准确率

研究人员为第二届 MLC-SLM 挑战赛开发了一个新颖的 speech-LLM 系统,专注于自动语音识别和说话人日志。他们的系统结合了 DiariZen-Large-s80 分割、CAM++ 说话人聚类以及 LoRA 适配的 omniASR LLM 7B v2 识别器,在开发集上取得了 29.27% 的宏观 tcpMER,显著优于官方基线。研究还分析了工程选择的影响,发现基于嵌入的说话人聚类比端到端方法更有效,并且重叠感知分割可能无意中增加 tcpMER。 AI

影响 该系统展示了语音识别和说话人日志方面的进步,可能提高多语言对话式 AI 的性能。

排序理由 该集群包含一篇详细介绍学术挑战新颖系统的研究论文。

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

Speech-LLM 系统在 MLC-SLM 挑战赛中取得高准确率

报道来源 [2]

  1. arXiv cs.AI TIER_1 English(EN) · Shuming Fang, Shuifei Zeng ·

    基于全语言ASR的Speech-LLM系统,用于第二届MLC-SLM挑战赛

    arXiv:2607.12468v1 Announce Type: cross Abstract: We describe our submission to Task 1 of the 2nd MLCSLM Challenge: a cascaded diarization-then-recognition system that combines DiariZen-Large-s80 (WavLM-Large) segmentation, CAM++ embedding-based two-speaker clustering, and a LoRA…

  2. arXiv cs.AI TIER_1 English(EN) · Shuifei Zeng ·

    基于全语言ASR的Speech-LLM系统,用于第二届MLC-SLM挑战赛

    We describe our submission to Task 1 of the 2nd MLCSLM Challenge: a cascaded diarization-then-recognition system that combines DiariZen-Large-s80 (WavLM-Large) segmentation, CAM++ embedding-based two-speaker clustering, and a LoRA-adapted omniASR LLM 7B v2 recognizer, with no ora…