Researchers have developed a novel framework to detect pathological hallucinations in Sinhala-to-English neural machine translation. This system utilizes a synthetic dataset of 45,000 samples, created using five corruption strategies and a semantic rescue mechanism. The fine-tuned mDeBERTa-v3 model achieved a token-level F1 score of 0.841, and an ensemble of neural risk scores, sequence log-probabilities, and cross-lingual embeddings further enhanced detection accuracy, with an AUROC of 0.970. The study also revealed significant variations in hallucination rates across eight different NMT systems. AI
IMPACT This research could lead to more reliable machine translation systems, particularly for low-resource languages, by improving the detection and mitigation of translation errors.
RANK_REASON The cluster contains an academic paper detailing a new method for hallucination detection in machine translation. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →