A new benchmark called NoteVQA has been developed to evaluate vision-language models (VLMs) on real-world visual questions, addressing the limitations of existing benchmarks that focus on predefined capabilities. The benchmark, curated from questions on the Chinese image-sharing platform Xiaohongshu, includes 252 items across 12 categories and 7 user intents, featuring human-audited interleaved answers with visual evidence. Evaluations of ten frontier VLMs showed that even with agentic search, the highest short-answer accuracy reached only 52.8%, and interleaved answer quality lagged behind human references, highlighting significant challenges for current VLMs in handling diverse, everyday visual queries. AI
IMPACT Highlights limitations in current VLMs for real-world applications, suggesting a need for improved visual reasoning and explanation capabilities.
RANK_REASON New academic paper introducing a novel benchmark for evaluating AI models. [lever_c_demoted from research: ic=1 ai=1.0]
- AgenticInterleave
- arXiv
- Hugging Face
- IVR-12
- NoteVQA
- Qwen3.5-397B-A17B
- React
- Vision--Language Models
- Xiaohongshu
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →