Researchers have introduced InstanceBench, a new diagnostic benchmark designed to evaluate referential reasoning and target identity in referring expression segmentation (RES) models. This benchmark includes over 6,000 images and 25,000 human-verified expressions, with a focus on distinguishing between different types of referential logic and separating errors in target selection from mask generation. Initial evaluations on 22 RES models revealed that while the top model achieved 67.1% mIoU, its performance on identity-aware metrics was lower, highlighting target selection as a primary bottleneck. AI
IMPACT This benchmark could lead to more robust AI models capable of better understanding and reasoning about object identity and relationships in images.
RANK_REASON The item describes a new academic benchmark and evaluation of AI models. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- Connected Papers
- CORE Recommender
- DagsHub
- Gotit.pub
- Hugging Face
- InstanceBench
- Litmaps
- ScienceCast
- scite Smart Citations
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →