A new research paper published on arXiv investigates gradient-based data attribution methods commonly used in large language models. The study reveals that these methods primarily track answer format similarity rather than task semantics. Researchers demonstrated this by creating benchmarks with shared tasks but varied answer formats, and vice versa, finding strong alignment only when answer formats matched. The findings suggest that current gradient-based attribution methods may not reliably identify task-relevant skills and could be influenced by superficial similarities in data. AI
影响 Challenges the reliability of current methods for analyzing and selecting training data for LLMs, potentially impacting future model development and evaluation.
排序理由 Academic paper published on arXiv detailing new research findings. [lever_c_demoted from research: ic=1 ai=1.0]
AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →