Researchers have developed a novel phoneme-guided initialization method to improve large language model (LLM) performance in automatic speech recognition (ASR), particularly in low-resource scenarios. This approach involves pre-training the audio encoder on speech-to-phoneme (S2P) tasks and the LLM on phoneme-to-grapheme (P2G) tasks before end-to-end fine-tuning. Experiments across multiple languages, including Japanese, Chinese, Tatar, and Urdu, demonstrated that this method matches or surpasses existing cascaded and end-to-end ASR models. Additionally, advancements in multilingual LLM-based P2G for ASR have been made, focusing on robustness strategies to handle S2P uncertainty and data imbalance, leading to reduced word error rates on benchmarks like CV-Lang10. AI
IMPACT Improves LLM performance in low-resource speech recognition and advances multilingual P2G capabilities.
RANK_REASON Two arXiv papers detailing novel methods for improving LLM-based speech recognition.
- arXiv
- CV-Lang10
- Hugging Face
- Japanese
- LLM
- phoneme-guided initialization
- phoneme-to-grapheme (P2G)
- speech-to-phoneme (S2P)
- Chinese
- Tatar
- Urdu
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →