Researchers have introduced MCR-Bench, a novel benchmark designed to evaluate the capabilities of large language models (LLMs) in realistic, multi-round code review scenarios. Unlike previous static approaches, MCR-Bench captures the dynamic, iterative nature of code review, incorporating defect metadata and cross-round state annotations across five programming languages. Experiments reveal that current mainstream LLMs struggle with defect detection and state tracking, particularly as the number of interaction rounds increases, and show varying performance across different defect types and severity levels. AI
IMPACT This benchmark could drive improvements in LLM capabilities for complex, interactive software development tasks.
RANK_REASON The item describes a new benchmark for evaluating LLMs in a specific research area (code review). [lever_c_demoted from research: ic=1 ai=1.0]
- code review
- Defect detection in lyophilized drug products with convolutional neural networks
- large language models
- machine language
- MCR-Bench
- software development
- type I and type II errors
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →