PulseAugur
实时 10:57:15
English(EN) PrefSQA: Pairwise Preference Prediction for Speech Quality Assessment and the Critical Role of High Quality Datasets

研究揭示音频大模型对评估差异敏感

一篇新发表在arXiv上的研究论文,考察了大型音频语言模型(LALMs)在使用多项选择问答(MCQA)框架进行评估时的鲁棒性。研究人员发现,像Audio Flamingo 2、Audio Flamingo 3、Qwen2.5-Omni-7B-Instruct和Kimi-Audio-7B-Instruct等模型对问题和选项措辞的变化以及选项呈现顺序很敏感。为解决这些局限性,该研究提出了一种更鲁棒的LALMs评估协议和指标。 AI

影响 强调了对音频语言模型需要更严格和标准化的评估方法,以确保可靠的性能评估。

排序理由 该集群包含讨论音频语言模型和语音质量评估方法的学术论文。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

研究揭示音频大模型对评估差异敏感

报道来源 [2]

  1. arXiv cs.CL TIER_1 English(EN) · Fernando L\'opez, Santosh Kesiraju, Jordi Luque ·

    大型音频语言模型在多项选择评估中的鲁棒性评估

    arXiv:2510.04584v2 Announce Type: replace Abstract: Recent advances in large audio language models (LALMs) have primarily been assessed using a multiple-choice question answering (MCQA) framework. However, subtle changes, such as shifting the order of choices, result in substanti…

  2. Hugging Face Daily Papers TIER_1 English(EN) ·

    PrefSQA:语音质量评估的成对偏好预测及其高质量数据集的关键作用

    Mean opinion scores (MOS) are widely used for speech quality assessment, yet scalar labels are sensitive to rater variability and listening test differences. This introduces labeling noise, which limits the reliability of MOS prediction. Preference prediction reduces this variabi…