A new study published on arXiv investigates the effectiveness of Large Language Model (LLM)-based Automated Program Repair (APR) techniques. The research analyzed how factors like bug complexity, fault localization accuracy, and the cost of different LLMs impact repair performance. Findings indicate that while complex bugs and imprecise fault localization pose challenges, LLM-based APR remains competitive, though higher-cost models do not always offer better cost-efficiency. Specifically, GPT-5 outperformed DeepSeek V4-Pro and DeepSeek V3.2 in repairing complex bugs, but DeepSeek V3.2 demonstrated superior cost-efficiency. AI
IMPACT Provides insights into optimizing LLM usage for software development tasks, highlighting cost-efficiency trade-offs.
RANK_REASON Research paper analyzing LLM performance on automated program repair. [lever_c_demoted from research: ic=1 ai=1.0]
- ChatRepair
- CodeCorrector
- DeepSeek
- DeepSeek V3.2
- DeepSeek V4-Pro
- generative pre-trained transformer
- GPT-5
- llama
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →