Researchers have introduced RefBench-PRO, a new benchmark designed to evaluate the perceptual and reasoning capabilities of Multi-modal Large Language Models (MLLMs) in Referring Expression Comprehension (REC). This benchmark decomposes REC into perception and reasoning dimensions across six challenging tasks, addressing the limitations of existing benchmarks that primarily focus on perceptual abilities. A novel automated data-generation pipeline and an RL-based learning scheme called Ref-R1 were also developed to enhance localization accuracy and provide a stronger baseline for REC. AI
IMPACT This benchmark could lead to more robust evaluation of multi-modal models, driving progress in visual-language understanding.
RANK_REASON The item is an academic paper detailing a new benchmark and methodology for evaluating AI models. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- Dynamic IoU-based GRPO
- Gotit.pub
- Hao Li
- Hugging Face
- Multi-modal Large Language Model
- RefBench-PRO
- Referring Expression Comprehension
- Ref-R1
- ScienceCast
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →