Researchers have introduced MT-Web2Code, a novel benchmark designed to evaluate Large Vision-Language Models (LVLMs) on complex, multi-turn coding tasks. This benchmark addresses the limitations of existing tools by focusing on iterative processes like reconstructing missing regions and modifying localized elements within web interfaces, rather than just full-page generation. MT-Web2Code includes 102 tasks across 16 domains and utilizes a Reverse-Corruption Trajectory Engine to create defects for deterministic evaluation. Initial experiments show that current coding agents struggle with these nuanced tasks, particularly in maintaining code integrity across multiple turns. AI
IMPACT This benchmark could drive advancements in AI agents capable of more complex and iterative software development tasks.
RANK_REASON The item describes a new benchmark and evaluation protocol for AI models, published on arXiv. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- CORE Recommender
- DagsHub
- Gotit.pub
- Hugging Face
- Large Vision-Language Models
- MT-Web2Code
- Reverse-Corruption Trajectory Engine
- ScienceCast
- VLM-based rubric
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →