A new study published on arXiv examines the robustness of large audio language models (LALMs) when evaluated using multiple-choice question answering (MCQA) frameworks. Researchers found that models like Audio Flamingo 2, Audio Flamingo 3, Qwen2.5-Omni-7B-Instruct, and Kimi-Audio-7B-Instruct are sensitive to variations in question and choice phrasing, as well as the order of presented options. To address these limitations, the study proposes a more robust evaluation protocol and metric for LALMs. AI
IMPACT Highlights the need for more rigorous and standardized evaluation methods for audio language models to ensure reliable performance assessment.
RANK_REASON The cluster contains academic papers discussing evaluation methodologies for audio language models and speech quality assessment.
Read on Hugging Face Daily Papers →
- Hugging Face
- Mean opinion scores
- PrefSQA
- arXiv
- Audio Flamingo 2
- Audio Flamingo 3
- Kimi-Audio-7B-Instruct
- Qwen2.5-Omni-7B-Instruct
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →