Researchers are exploring new methods for multilingual Automatic Speech Recognition (ASR), particularly for code-switching scenarios where multiple languages are used within a single conversation. One paper investigates generalizing code-switching capabilities to unseen language pairs through model merging, finding limited success. Another project, BaltiVoice, introduces a new speech corpus and fine-tuned Whisper model for the Balti language, significantly improving ASR accuracy. Additionally, a system called WAXAL-NET demonstrates that specialized, smaller ASR models can outperform large multilingual models for African languages, and a real-time multilingual ASR system uses a routing approach with smaller, specialized models to achieve high accuracy and efficiency. AI
IMPACT Advances in multilingual ASR could significantly improve human-AI interaction across diverse linguistic communities and enable more efficient, specialized speech recognition systems.
RANK_REASON Multiple research papers and projects presenting new models, datasets, and techniques for Automatic Speech Recognition (ASR), particularly in multilingual and code-switching contexts.
- Gladia
- JeanMichelRanu
- Silero VAD
- SpeechBrain
- Zipformer
- WAXAL-NET
- BaltiVoice
- OpenAI
- WAXAL
- Whisper
- Hugging Face
- Mozilla Common Voice
AI-generated summary · Google Gemini · from 9 sources. How we write summaries →