PulseAugur
EN
LIVE 16:04:27

CPU TTS benchmark: Kokoro 82M leads in quality, Inflect-Nano-v1 in speed

A benchmark comparing three open-weight Text-to-Speech (TTS) models—Kokoro 82M, Supertonic 3, and Inflect-Nano-v1—on a CPU revealed significant performance and quality differences. Inflect-Nano-v1, despite its small parameter count and fastest real-time factor (RTF) of 0.1376, was found to be over-rated by UTMOS scoring and suffers from a hard output length limitation. Supertonic 3 offered a trade-off, with a 5-step configuration achieving a MOS of 4.37 at an RTF of 0.3164, while Kokoro 82M, though the slowest with RTFs between 0.5711 and 0.7865, produced the most human-like audio. AI

IMPACT Provides insights into the trade-offs between speed and audio quality for CPU-based TTS models, guiding developers on model selection.

RANK_REASON The cluster details a benchmark comparing multiple open-weight TTS models, including performance metrics and quality assessments. [lever_c_demoted from research: ic=1 ai=1.0]

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

CPU TTS benchmark: Kokoro 82M leads in quality, Inflect-Nano-v1 in speed

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
The cluster details a benchmark comparing multiple open-weight TTS models, including performance metrics and quality assessments. [lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
95 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

Full methodology in our editorial standards.

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/gvij ·

    CPU-only TTS benchmark: Kokoro 82M vs Supertonic 3 vs Inflect-Nano-v1 (4.6M params), with UTMOS scoring on every sample

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1udg3rf/cpuonly_tts_benchmark_kokoro_82m_vs_supertonic_3/"> <img alt="CPU-only TTS benchmark: Kokoro 82M vs Supertonic 3 vs Inflect-Nano-v1 (4.6M params), with UTMOS scoring on every sample" src="https://previ…