Researchers have developed SPRINT, a new benchmark designed to evaluate the proactive risk inference capabilities of multimodal large language models (MLLMs). The benchmark utilizes 2,888 real-world sports videos, including detailed annotations of hazard cues and accident causes, to test MLLMs' ability to predict physical dangers. Current state-of-the-art MLLMs show a significant gap, excelling at hazard detection but struggling to identify the underlying causes, and exhibit a tendency to generate false alarms on safe videos. AI
IMPACT Highlights limitations in current MLLMs for real-world physical safety applications, indicating a need for improved causal reasoning.
RANK_REASON The item describes a new benchmark and research paper published on arXiv. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →