Researchers have investigated the alignment of reinforcement learning rewards with human perception in codec-based text-to-speech (TTS) models. Using Group Relative Policy Optimization (GRPO) with subjective rewards for style, naturalness, and likability, they found that each reward primarily improved its specific target metric, indicating that subjective predictors are not interchangeable quality surrogates. Human A/B tests showed uneven transfer of these rewards, and a reward-gap analysis suggested that while signed reward gaps predict listener choices, per-axis calibration remains heterogeneous. The study also found that a Best-of-8 reranking approach served as a strong human-level baseline, comparable to GRPO in perceptual quality. AI
IMPACT This research highlights the challenges in aligning AI-generated speech with human preferences, suggesting a need for more nuanced reward mechanisms in TTS model training.
RANK_REASON Academic paper on AI model training methodology. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →