Researchers are exploring new methods for assessing speech quality beyond traditional Mean Opinion Scores (MOS). One paper introduces PrefSQA, which uses pairwise preference prediction to reduce rater variability and improve reliability, particularly with high-quality preference datasets. Another study investigates discrepancies between human listeners and MOS prediction models, finding that while models track acoustic degradation, they often miss prosodic errors and exhibit biases in speaker characteristics. A third paper proposes NVMOS for assessing the quality of non-verbal vocalizations, demonstrating that current multimodal large language models like Gemini struggle with this task and do not reliably replace human expert judgment. AI
IMPACT Advances in speech quality assessment could improve the development and evaluation of text-to-speech systems and other audio technologies.
RANK_REASON Multiple research papers published on arXiv detailing new methods and findings in speech quality assessment.
- arXiv
- Gemini
- NVMOS
- NV-TTS
- acoustic degradation
- fundamental frequency
- Mean opinion score
- MOS prediction models
- Pitch
- prosodic errors
- speaker characteristics
- speech synthesis
- PrefSQA
AI-generated summary · Google Gemini · from 4 sources. How we write summaries →