FermionResearch has released Phonon-2, an open-source speech recognition model that achieves high accuracy for English audio with a small file size. The model averages 5.21% word error rate on the Open ASR Leaderboard, outperforming larger models and matching the performance of its larger, full-precision teacher model while being 15 times smaller. Phonon-2 is designed for efficient transcription, capable of processing an hour of audio in approximately 20 seconds on an Apple M5 MacBook Air and significantly faster on high-end hardware like NVIDIA H100 GPUs. AI
IMPACT Provides a highly accurate and efficient open-source option for speech-to-text applications, potentially lowering barriers for developers.
RANK_REASON Release of an open-source model with benchmark performance data. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Hugging Face Trending Models →
- Apple Inc.
- Atomic Chat
- Deta
- Docker
- FermionResearch/Phonon-2
- Google Colab
- Kaggle
- Linux
- LM Studio
- Microsoft Windows
- Mlx
- Nvidia
- NVIDIA Parakeet-TDT 0.6B v3
- Open ASR Leaderboard
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →