A new study published on arXiv investigates the impact of data contamination on code intelligence models, specifically examining how different types of contamination affect performance evaluations. The research tested various models including RoBERTa, GPT-2, LLaMA, and StarCoder across code translation, generation, and summarization tasks in Java and Python. Findings suggest that while paired contamination does not significantly overestimate performance in pre-training/fine-tuning paradigms, it does affect LLMs during direct inference or small-scale fine-tuning. AI
IMPACT Provides new insights into the evaluation and deployment of code intelligence models, challenging conventional beliefs about data contamination.
RANK_REASON Research paper published on arXiv detailing empirical study of data contamination in code intelligence models. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →