Researchers have identified a critical boundary within Conformer-Large automatic speech recognition (ASR) models where hallucinations, or fluent text unrelated to the audio, can emerge. By studying two independently trained models (one CTC and one RNN-T) under degraded conditions, they found that bypassing the final encoder stage consistently led to divergence and the potential for hallucination. This stage is where representations become more compact and text becomes readable by the decoder, suggesting a mechanistic precondition for grounded recognition failure. AI
IMPACT Identifies a specific failure mode in ASR models, potentially leading to improved robustness and accuracy in real-world applications.
RANK_REASON The cluster contains a research paper detailing findings about ASR model behavior. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- Conformer-Large
- DagsHub
- Gotit.pub
- Hugging Face
- Influence Flower
- RNN-T
- ScienceCast
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →