PulseAugur
EN
LIVE 02:03:38

audio.cpp 0.5 adds DramaBox TTS, Confucius4 voice transfer, and AMD GPU support

The audio.cpp project has released version 0.5, introducing several new AI models and improved hardware support. Key additions include DramaBox, an expressive TTS model capable of controlling emotion and delivery, and Confucius4, which enables cross-lingual voice transfer. The release also features RVC for voice conversion, BS-RoFormer for vocal separation, and integrates models like GLM-TTS, Kroko ASR, Parakeet-TDT, Inflect Micro v2, and Fun-ASR-Nano. Performance enhancements include faster Metal support on Apple Silicon and early ROCm/HIP support for AMD GPUs, alongside improvements to live PCM ingest and streaming transcript deltas. AI

IMPACT Expands the capabilities of open-source audio processing tools with new expressive and cross-lingual voice models.

RANK_REASON Release of a new version of an open-source audio processing project with new models and hardware support.

Read on r/StableDiffusion →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

audio.cpp 0.5 adds DramaBox TTS, Confucius4 voice transfer, and AMD GPU support

COVERAGE [1]

  1. r/StableDiffusion TIER_2 English(EN) · /u/Acceptable-Cycle4645 ·

    [audio.cpp] Release 0.5: DramaBox expressive TTS, Confucius4 cross-lingual voice transfer, plus 7 more models and ROCm/HIP

    <table> <tr><td> <a href="https://www.reddit.com/r/StableDiffusion/comments/1vc8nh6/audiocpp_release_05_dramabox_expressive_tts/"> <img alt="[audio.cpp] Release 0.5: DramaBox expressive TTS, Confucius4 cross-lingual voice transfer, plus 7 more models and ROCm/HIP" src="https://ex…