PulseAugur
实时 08:40:16
English(EN) Language-Aware Distillation for Multilingual Instruction-Following Speech LLMs with ASR-Only Supervision

新的MEUSLI投影仪支持多语言ASR和语音理解

研究人员开发了MEUSLI,这是一种新颖的多语言投影仪,旨在将语音编码器与大语言模型(LLM)连接起来,以完成高级语音处理任务。该系统扩展了现有的单语投影仪,支持28种欧洲语言的端到端自动语音识别(ASR),并且可以进一步适应其他语言。MEUSLI还展示了超越ASR的能力,只需极少的特定任务监督即可实现多语言语音翻译和主题识别。 AI

影响 推进了多语言语音理解和翻译能力,有可能拓宽基于LLM的语音技术的获取途径。

排序理由 该集群描述了详细介绍使用LLM进行多语言语音处理的新方法和系统的新研究论文。

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

新的MEUSLI投影仪支持多语言ASR和语音理解

报道来源 [2]

  1. arXiv cs.CL TIER_1 English(EN) · Lorenzo Concina, Seraphina Fong, Marco Matassoni, Alessio Brutti ·

    MEUSLI: a Multilingual Projector for LLM-based ASR and Beyond

    arXiv:2607.22100v1 Announce Type: new Abstract: Lightweight projectors are an established way to connect pre-trained speech encoders with large language models (LLMs), mapping acoustic features into token-level embeddings for tasks like ASR and spoken question answering. Existing…

  2. arXiv cs.CL TIER_1 English(EN) · Shreyas Gopal, Donghang Wu, Ashutosh Anshul, Yeo Yue Heng, Yizhou Peng, Haoyang Li, Hexin Liu, Eng Siong Chng ·

    Language-Aware Distillation for Multilingual Instruction-Following Speech LLMs with ASR-Only Supervision

    arXiv:2603.07025v2 Announce Type: replace Abstract: Speech Large Language Models (LLMs) that understand and follow instructions in many languages are useful for real-world interaction, but are difficult to train with supervised fine-tuning, requiring large, task-specific speech c…