The audio.cpp project has released version 0.5, introducing several new AI models and improved hardware support. Key additions include DramaBox, an expressive TTS model capable of controlling emotion and delivery, and Confucius4, which enables cross-lingual voice transfer. The release also features RVC for voice conversion, BS-RoFormer for vocal separation, and integrates models like GLM-TTS, Kroko ASR, Parakeet-TDT, Inflect Micro v2, and Fun-ASR-Nano. Performance enhancements include faster Metal support on Apple Silicon and early ROCm/HIP support for AMD GPUs, alongside improvements to live PCM ingest and streaming transcript deltas. AI
IMPACT Expands the capabilities of open-source audio processing tools with new expressive and cross-lingual voice models.
RANK_REASON Release of a new version of an open-source audio processing project with new models and hardware support.
- Apple Silicon
- audio.cpp
- BS-RoFormer
- Confucius4
- DramaBox
- FunASR
- Fun-ASR-Nano
- GLM-TTS
- Inflect Micro v2
- Kroko ASR
- Parakeet-TDT
- ROCm
- StableDiffusion
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →