A new research paper identifies "deletion avoidance" as a key issue in large language models' code editing capabilities, where models tend to retain code that should be removed. This behavior, observed across leading models on the SWE-bench benchmark, leads to codebases that are harder to maintain. The study introduces the CanItDelete benchmark to specifically test deletion tasks and finds that even advanced models struggle, often resorting to "Guard-and-Go" patterns instead of outright removal. Promisingly, the research suggests that training models with deletion-focused data can mitigate this issue and improve overall code-editing performance. AI
IMPACT Highlights a critical flaw in LLM code editing that could hinder enterprise adoption and requires further training to address.
RANK_REASON The cluster reports on a new academic paper detailing research findings and a new benchmark for evaluating LLM code editing capabilities.
- Amir M. Ebrahimi
- CanItDelete
- GPT 5.6 "Sol"
- Guard-and-Go
- large-language models
- SWE-bench
- Claude (Opus 4.8)
- DeepSeek V4-Pro
- GLM-5.2
- Hugging Face
- Kimi K2 Thinking
- Minimax M3
- Qwen3 235B
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →