PulseAugur
EN
LIVE 22:41:47

audio.cpp adds music, SFX, and long-form TTS models via C++/GGML

The audio.cpp project has released significant updates, introducing native C++/GGML support for several new audio models including ACE-Step, Stable Audio, HeartMuLa, RoFormer, and HTDemucs. This expansion enables music and SFX generation, as well as source separation, with some models now capable of generating up to 10 minutes of audio. Additionally, VibeVoice 1.5B has been integrated, demonstrating impressive performance for long-form TTS, generating 90 minutes of audio in under 23 minutes on an RTX 5090, significantly outperforming Python-based alternatives. AI

IMPACT Expands local audio model capabilities, offering faster inference for music, SFX, and long-form TTS compared to Python.

RANK_REASON This is a release of new models and features for a specific software project, audio.cpp, rather than a core AI model release from a major lab.

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 4 sources. How we write summaries →

audio.cpp adds music, SFX, and long-form TTS models via C++/GGML

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
This is a release of new models and features for a specific software project, audio.cpp, rather than a core AI model release from a major lab.
Source corroboration
4 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
model release, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
99 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

Full methodology in our editorial standards.

COVERAGE [4]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/Acceptable-Cycle4645 ·

    [audio.cpp] The Sound of GGML — C++/GGML native ACE-Step, Stable Audio, HeartMuLa, RoFormer, HTDemucs released. 10-Minute Music in 60 Seconds!

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1um2tbf/audiocpp_the_sound_of_ggml_cggml_native_acestep/"> <img alt="[audio.cpp] The Sound of GGML — C++/GGML native ACE-Step, Stable Audio, HeartMuLa, RoFormer, HTDemucs released. 10-Minute Music in 60 Second…

  2. r/LocalLLaMA TIER_1 English(EN) · /u/Acceptable-Cycle4645 ·

    [audio.cpp] VibeVoice 1.5B released — 90-min podcast in 22.95 min, 4.08x real-time, 2.86x faster than Python without quantization. Native C++/ggml

    <!-- SC_OFF --><div class="md"><p>I’m the author of audio.cpp, a C++/ggml runtime for local audio models.</p> <p>I just added VibeVoice 1.5B support and wanted to share the benchmark because long-form multi-speaker TTS is a good stress test for local inference runtimes.</p> <p>Re…

  3. r/StableDiffusion TIER_2 English(EN) · /u/Acceptable-Cycle4645 ·

    [audio.cpp] The Sound of GGML — C++/GGML native ACE-Step, Stable Audio, HeartMuLa, RoFormer, HTDemucs released. 10-Minute Music in 60 Seconds!

    <!-- SC_OFF --><div class="md"><p>Hi music fans, I just released a big music/audio expansion in <code>audio.cpp</code>.</p> <p>This batch adds <strong>music generation</strong>, <strong>SFX generation</strong>, and <strong>source separation</strong> to the released framework surf…

  4. r/StableDiffusion TIER_2 English(EN) · /u/Acceptable-Cycle4645 ·

    Native C++/ggml VibeVoice 1.5B released — 90-min podcast in 22.95 min, 4.08x real-time, 2.86x faster than Python without quantization.

    <!-- SC_OFF --><div class="md"><p>I’m the author of audio.cpp, a C++/ggml runtime for local audio models.</p> <p>I just added VibeVoice 1.5B support and wanted to share the benchmark because long-form multi-speaker TTS is a good stress test for local inference runtimes.</p> <p>Re…