Microsoft has released VibeVoice-ASR-BitNet, a highly compressed automatic speech recognition model designed for real-time inference on edge CPUs without requiring a GPU. This model achieves significant speedups over existing solutions like Whisper.cpp, with a real-time factor (RTF) below 1, using as few as three CPU threads. The compression is achieved through heterogeneous quantization, reducing the model size from 4.62 GB to 1.58 GB while maintaining comparable accuracy and supporting multiple languages. AI
IMPACT Enables real-time speech recognition on resource-constrained edge devices, potentially accelerating adoption in areas like IoT and mobile applications.
RANK_REASON Model release from a major tech company (Microsoft) with a focus on performance and efficiency for edge devices. [lever_c_demoted from frontier_release: ic=2 ai=1.0]
- Arm Holdings
- arXiv
- BITNET
- GGML
- Hugging Face
- variational auto-encoder
- VibeVoice-ASR
- VibeVoice-ASR-BitNet
- x86
- llama.cpp
- llama-cpp-python
- microsoft/VibeVoice-ASR-BitNet
- Ollama
AI-generated summary · Google Gemini · from 3 sources. How we write summaries →