Researchers have developed HealBench, a new benchmark designed to evaluate the effectiveness and safety of Large Language Models (LLMs) in automatically healing runtime errors within real-world software repositories. This system includes HealGuard, a safety mechanism that uses static and dynamic taint analysis to ensure that code generated by LLMs does not compromise protected operations. In evaluations, the best-performing LLM setting successfully resumed execution in 38.11% of instances and passed target tests in 28.68% of cases, though HealGuard flagged a significant portion of these as potentially unsafe. AI
IMPACT This research advances the safety and reliability of LLMs for automated code repair, potentially enabling more robust software development tools.
RANK_REASON The item is a research paper introducing a new benchmark and safety mechanism for LLM-based software error healing. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- Gotit.pub
- HealBench
- HealCore
- HealGuard
- Hugging Face
- Python
- ScienceCast
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →