PulseAugur
EN
LIVE 11:24:06

New research explores advanced speech quality assessment methods beyond MOS

Researchers are exploring new methods for assessing speech quality beyond traditional Mean Opinion Scores (MOS). One paper introduces PrefSQA, which uses pairwise preference prediction to reduce rater variability and improve reliability, particularly with high-quality preference datasets. Another study investigates discrepancies between human listeners and MOS prediction models, finding that while models track acoustic degradation, they often miss prosodic errors and exhibit biases in speaker characteristics. A third paper proposes NVMOS for assessing the quality of non-verbal vocalizations, demonstrating that current multimodal large language models like Gemini struggle with this task and do not reliably replace human expert judgment. AI

IMPACT Advances in speech quality assessment could improve the development and evaluation of text-to-speech systems and other audio technologies.

RANK_REASON Multiple research papers published on arXiv detailing new methods and findings in speech quality assessment.

Read on arXiv cs.CL →

AI-generated summary · Google Gemini · from 4 sources. How we write summaries →

New research explores advanced speech quality assessment methods beyond MOS

COVERAGE [4]

  1. arXiv cs.AI TIER_1 English(EN) · Junyi Fan, Donald S. Williamson ·

    PrefSQA: Pairwise Preference Prediction for Speech Quality Assessment and the Critical Role of High Quality Datasets

    arXiv:2606.19597v1 Announce Type: cross Abstract: Mean opinion scores (MOS) are widely used for speech quality assessment, yet scalar labels are sensitive to rater variability and listening test differences. This introduces labeling noise, which limits the reliability of MOS pred…

  2. arXiv cs.CL TIER_1 English(EN) · Masato Takagi, Masaya Kawamura, Reo Shimizu, Yuma Shirahata ·

    Investigating Human-Model Discrepancies in Speech Quality Assessment via Acoustic and Prosodic Perturbations

    arXiv:2606.19951v1 Announce Type: cross Abstract: Mean opinion score (MOS) prediction models are widely used as proxy metrics in text-to-speech (TTS) research, yet their ability to capture quality differences beyond acoustic fidelity remains unclear. We investigate this via contr…

  3. arXiv cs.CL TIER_1 English(EN) · Yuma Shirahata ·

    Investigating Human-Model Discrepancies in Speech Quality Assessment via Acoustic and Prosodic Perturbations

    Mean opinion score (MOS) prediction models are widely used as proxy metrics in text-to-speech (TTS) research, yet their ability to capture quality differences beyond acoustic fidelity remains unclear. We investigate this via controlled perturbations on speech: acoustic degradatio…

  4. arXiv cs.AI TIER_1 English(EN) · Jialong Mai, Jinxin Ji, Xiaofen Xing, Wencui Liu, Xiangmin Xu ·

    NVMOS: Non-Verbal Vocalization Quality Assessment in Speech

    arXiv:2606.15888v1 Announce Type: cross Abstract: Non-verbal vocalizations (NVs), such as laughter, sighs, and coughs, are important acoustic cues for emotion and intent. Existing speech quality assessment methods typically focus on overall naturalness, while non-verbal TTS evalu…