Researchers from Meta, Stanford, Harvard, and UW have developed SWE-sweep, a new benchmark designed to evaluate large language models' ability to proactively identify and fix bugs in large codebases before they impact users. The benchmark uses real-world bugs and scores models based on their success in finding and resolving these issues within a given codebase. Early results indicate that while some models struggle, Luna xhigh shows cost-efficiency, and the researchers are seeking recommendations for additional open-weight models to include in future updates. AI
IMPACT This benchmark could drive development of more robust AI coding assistants capable of proactive error detection.
RANK_REASON New benchmark and paper released by academic researchers. [lever_c_demoted from research: ic=1 ai=1.0]
- Claude 3
- Code Llama
- DeepSeek Coder
- GPT-4
- Harvard
- Llama 2
- LMSYS Chatbot Arena
- Luna xhigh
- Meta
- Mistral AI
- Mixtral
- Phind
- Stanford
- StarCoder2
- SWE-sweep
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →