Researchers have developed SrDetection, a novel framework designed to identify data leakage in code large language models (Code LLMs). This self-referential approach generates variations of benchmark samples to detect when a model's performance is artificially inflated due to prior exposure to the benchmark data. SrDetection offers improvements in both gray-box and black-box settings, outperforming existing methods and revealing specific leakage patterns across various Code LLMs and benchmarks. AI
IMPACT This framework could lead to more reliable evaluations of code LLMs, ensuring that benchmark performance accurately reflects true capabilities rather than memorization.
RANK_REASON The cluster describes a research paper detailing a new framework for detecting data leakage in code LLMs.
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →