A new study benchmarks six multilingual pre-trained models for Nepali Automatic Speech Recognition (ASR) using a standardized fine-tuning protocol. The research found that Whisper-Large-v3-Turbo and IndicWav2Vec performed comparably, demonstrating that language-family proximity in pretraining can be as effective as sheer scale for Nepali ASR. The study also highlighted that CTC decoders offer significantly faster performance than autoregressive models like Whisper at similar accuracy levels, making them preferable for applications with latency constraints. This work provides the first standardized, multi-model, efficiency-aware reference numbers for Nepali ASR. AI
IMPACT Provides standardized benchmarks for Nepali ASR, guiding model selection for efficiency and accuracy.
RANK_REASON The cluster contains an academic paper detailing a comparative analysis and benchmark of AI models. [lever_c_demoted from research: ic=1 ai=1.0]
- Common Voice
- Conformer-Hi
- IndicWav2Vec
- MMS-1B
- OpenSLR SLR54
- Open Speech and Language Resources
- Whisper Large v3 Turbo
- Whisper-medium
- XLSR-53
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →