Researchers have introduced HeiCo-FOCUS, a new dataset designed to evaluate the long-context video understanding capabilities of vision-language models (VLMs). This dataset, derived from Heidelberg Colorectal surgeries, focuses on the challenge of tracking foreign objects within procedures that can last for hours. HeiCo-FOCUS includes 30,000 visual question answering pairs and employs a multi-track evaluation framework to progressively test models on object recognition, temporal grounding, aggregation, event understanding, and complex reasoning. Initial experiments with ten frontier VLMs revealed significant challenges, particularly in temporal grounding, indicating that current models are far from mastering this complex task. AI
IMPACT This dataset aims to push the development of VLMs capable of sustained temporal reasoning, crucial for applications beyond short-form content.
RANK_REASON The item describes a new dataset and evaluation framework for long-context video understanding, presented in an academic paper. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- CatalyzeX
- DagsHub
- Foreign Object Contextual Understanding in Surgery
- Gotit.pub
- HeiCo-FOCUS
- Heidelberg
- Heidelberg Colorectal surgeries
- Hugging Face
- ScienceCast
- vision-language model
- visual question answering
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →