A new research paper published on arXiv highlights significant overestimation in the evaluation of sign language translation (SLT) models. The study found that current evaluation methods, which often include overlapping signers across training and testing datasets, lead to inflated performance scores. When evaluated using signer-independent protocols, the performance of leading SLT models like GFSLT-VLP, GASLT, and SignCL dropped dramatically, with BLEU-4 scores falling from over 21 to as low as 3.59 on the PHOENIX14T dataset. The researchers recommend adopting signer-independent evaluation, restructuring datasets for sentence-disjoint splits, and reporting both dependent and independent results to ensure more accurate benchmarking and transparency in SLT capabilities. AI
IMPACT Highlights critical flaws in current SLT evaluation, potentially leading to more robust and generalizable models.
RANK_REASON Academic paper detailing a new evaluation methodology for sign language translation models. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →