The audio.cpp project has released version 0.5, introducing significant updates to its text-to-speech (TTS) and voice transfer capabilities. A key highlight is DramaBox, a new model designed for prompt-directed voice acting with control over emotion and delivery. The release also features Confucius4-TTS for cross-lingual voice transfer and integrates several other models and ASR systems. Enhancements to platform support include early HIP/ROCm compatibility for AMD GPUs and performance improvements for Apple Silicon, alongside better live streaming and ingest features. AI
IMPACT Enhances local audio inference capabilities with new expressive TTS and cross-lingual voice transfer models.
RANK_REASON This is a software release for a specific tool, not a frontier model release or significant industry event.
- Apple Silicon
- audio.cpp
- BS-RoFormer
- Confucius4
- DramaBox
- FunASR
- Fun-ASR-Nano
- GLM-TTS
- Heterogeneous Integration Platform
- Inflect Micro v2
- Kroko ASR
- Parakeet-TDT
- Rocm
- StableDiffusion
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →