A new study analyzing hallucination in Large Language Models (LLMs) used for automated program repair (APR) reveals significant issues. Researchers examined three LLMs across 832 Defects4J bugs, finding that only 21.0%-55.9% of generated patches passed developer-written test suites. The study identified repair hallucinations in 72.7% of analyzed cases, with incorrect localization and repair strategies being common causes. Additionally, LLMs frequently misidentified triggering test cases and inaccurately predicted line coverage. AI
IMPACT Highlights critical limitations in LLM-driven code repair, suggesting a need for improved hallucination detection and mitigation strategies.
RANK_REASON Academic paper detailing a study on LLM hallucination in automated program repair. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →