PulseAugur
实时 08:50:15
Română(RO) Xiaomi-CocktailASR-1 Technical Report

小米发布基于LLM的嘈杂环境下的ASR系统

研究人员推出了一种新颖的端到端自动语音识别(ASR)架构Xiaomi-CocktailASR-1,旨在解决多说话人环境中的“鸡尾酒会问题”。该基于LLM的系统利用声纹提示,无需预先分离即可转录目标说话人的语音,同时保持具有竞争力的单说话人性能,并增加了对缺席目标说话人的关键拒绝能力。该架构还支持链式思考(Chain-of-Thought)推理模式,并在各种基准测试中取得了最先进的结果。 AI

影响 该模型有望显著提高在嘈杂、多说话人环境下的ASR性能,造福于语音助手和转录服务等应用。

排序理由 该集群包含一篇在arXiv上发表的技术报告,详细介绍了用于自动语音识别的新模型架构。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

小米发布基于LLM的嘈杂环境下的ASR系统

本文如何被排名

Signal score
16 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群包含一篇在arXiv上发表的技术报告,详细介绍了用于自动语音识别的新模型架构。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
model release, paper
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. arXiv cs.CL TIER_1 Română(RO) · Yiru Zhang, Hang Su, Lichun Fan, Ying Zeng, Chang Liu, Yifeng Wang, Yuquan Liang, Tao Li, Lian Li, Wenhao Yang, Jian Luan, Cong Zou, Heng Qu ·

    Xiaomi-CocktailASR-1 技术报告

    arXiv:2609.11274v1 Announce Type: cross Abstract: Recently, large language model (LLM) based ASR models have achieved significant progress, yet they generally lack support for multi-speaker scenarios, where the cocktail party problem remains a critical bottleneck for further adva…