A new research paper published on arXiv addresses a significant bias in the COMET metric, which is used to evaluate machine translation quality. The study found that COMET scores are heavily influenced by the script used to write the target language, rather than solely reflecting translation accuracy. This script-induced bias accounts for a substantial portion of COMET's score variance and negatively impacts its agreement with human annotators across different Indic languages. The researchers propose a method called COMET-QN to normalize scores across scripts and offer diagnostic tools to identify and quantify this bias, advocating for greater transparency in reporting evaluation results. AI
IMPACT Highlights a critical flaw in a widely used MT evaluation metric, potentially impacting future research and development in machine translation, especially for multilingual contexts.
RANK_REASON Research paper published on arXiv detailing a new diagnostic and correction method for a machine translation evaluation metric. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →