A new research paper from arXiv explores the effectiveness of reference-free speech quality metrics, such as UTMOS, DNSMOS, and SCOREQ, in evaluating modern text-to-speech (TTS) systems. The study found that while these metrics can identify audible defects, they struggle to reliably distinguish listener preferences in high-quality audio. The research proposes a composite metric as a more robust evaluator and notes that optimizing TTS models with single-score rewards can lead to undesirable "reward hacking," where the metric improves but actual human perception of quality declines. AI
IMPACT Highlights limitations of current automated evaluation metrics for TTS, suggesting a need for more sophisticated approaches to ensure perceived audio quality.
RANK_REASON Research paper evaluating metrics for text-to-speech systems. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →