Researchers have introduced ExBind, a new diagnostic benchmark designed to evaluate the accuracy of multimodal models in mapping visual or semantic references to specific executable objects. This benchmark isolates the visual-to-executable correspondence layer, moving beyond simple execution success to pinpoint failures in object selection. ExBind includes a broad suite of 250 cases and a targeted suite of 240 cases, formatted in various structures like SVG, DOM, and tables. Early evaluations show Qwen2.5-VL-3B achieving 76.4% exact accuracy, while Qwen3-VL-4B reached 98.8% exact accuracy on the benchmark. AI
IMPACT This benchmark could drive improvements in multimodal models' ability to accurately interpret and interact with visual interfaces.
RANK_REASON The item describes a new benchmark for evaluating multimodal models, presented in an arXiv paper. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →