Researchers have developed a new benchmark to evaluate automated Text-to-Speech (TTS) systems, moving beyond simple naturalness metrics. The benchmark deconstructs speech quality into 10 distinct linguistic dimensions, using annotations from trained linguists on 860 utterances. Evaluations of four Mean Opinion Score (MOS) predictors and four Audio Large Language Models (Audio-LLM) judges showed that MOS predictors focus solely on acoustic signal quality, while Audio-LLMs exhibit prompt-dependent detection that doesn't cover all dimensions. Neither approach reliably identifies a broad range of linguistically structured speech errors, highlighting the need for more nuanced TTS evaluation. AI
IMPACT This research introduces a more nuanced evaluation framework for TTS models, potentially leading to more accurate and linguistically aware speech synthesis systems.
RANK_REASON The cluster contains a research paper detailing a new benchmark for evaluating TTS systems. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- Audio Large Language Models
- CatalyzeX
- DagsHub
- Gotit.pub
- Hugging Face
- Mean Opinion Score
- ScienceCast
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →