PulseAugur
EN
LIVE 12:19:43

Kyutai's Pocket TTS offers CPU-based voice cloning from 5s audio

Kyutai has released Pocket TTS, a ~100M parameter streaming language model that generates audio tokens autoregressively. This model is notable for its ability to perform zero-shot voice cloning from just 5 seconds of audio, a feature not present in other CPU-friendly models like Kokoro, Supertonic, and Inflect-Nano. While Pocket TTS is the slowest among the tested models, it offers voice cloning capabilities on CPU without requiring a GPU, making it a unique option for interactive applications. AI

IMPACT Enables voice cloning on consumer hardware, potentially lowering the barrier for AI-powered audio generation.

RANK_REASON New product release from a smaller developer, offering a specific niche capability (CPU voice cloning).

Read on r/StableDiffusion →

AI-generated summary · Google Gemini · from 2 sources. How we write summaries →

Kyutai's Pocket TTS offers CPU-based voice cloning from 5s audio

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
New product release from a smaller developer, offering a specific niche capability (CPU voice cloning).
Source corroboration
2 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
model release, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
90 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

Full methodology in our editorial standards.

COVERAGE [2]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/gvij ·

    Kyutai's Pocket TTS clones a voice from 5 seconds of audio, on CPU, under MIT. Benchmarked against Kokoro, Supertonic, and Inflect-Nano for Eng. TTS

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1up07mk/kyutais_pocket_tts_clones_a_voice_from_5_seconds/"> <img alt="Kyutai's Pocket TTS clones a voice from 5 seconds of audio, on CPU, under MIT. Benchmarked against Kokoro, Supertonic, and Inflect-Nano for…

  2. r/StableDiffusion TIER_2 English(EN) · /u/gvij ·

    Voice cloning on CPU, no GPU needed. Benchmarked 4 open TTS models including Kyutai's new Pocket TTS

    <table> <tr><td> <a href="https://www.reddit.com/r/StableDiffusion/comments/1up0my7/voice_cloning_on_cpu_no_gpu_needed_benchmarked_4/"> <img alt="Voice cloning on CPU, no GPU needed. Benchmarked 4 open TTS models including Kyutai's new Pocket TTS" src="https://preview.redd.it/w6e…