Researchers have developed a new framework to evaluate the robustness of low-resource multilingual text-to-speech (TTS) systems when processing complex text inputs. This framework assesses content consistency, language consistency, and generation stability across languages like Thai, Vietnamese, Swahili, and Indonesian. The study introduces automatic diagnostic metrics and a Text Risk Score (TRS) to predict synthesis risks from text features without requiring manual annotation or model training. Experiments on OmniVoice, VoxCPM2, and MMS-TTS revealed distinct failure patterns in handling numbers, dates, named entities, and code-switched expressions, highlighting the limitations of current evaluation methods. AI
IMPACT This research could lead to more reliable and robust multilingual TTS systems by identifying and mitigating failure points in complex text processing.
RANK_REASON The cluster contains an academic paper detailing a new evaluation framework and metrics for TTS systems. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →