A practical benchmark evaluated four voice cloning models (Pocket TTS, Kokoro, Audio88 & Yassin, and XTTS v2) on CPU performance. The evaluation focused on speaker similarity, naturalness, intelligibility, and latency, using the VCTK dataset with diverse speakers and accents. The study found that while XTTS v2 and Audio88 & Yassin are true zero-shot cloners, Kokoro performed better in naturalness, and Pocket TTS and Kokoro served as useful baselines for CPU quality and speed. The benchmark was conducted by an AI agent named Neo, which also identified and fixed a data bug in the VCTK dataset. AI
IMPACT Provides practical insights for developers choosing voice cloning models for CPU-based applications, highlighting trade-offs between cloning fidelity, naturalness, and speed.
RANK_REASON The cluster describes a practical benchmark and evaluation of existing voice cloning models, including methodology and dataset analysis.
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →