Researchers have introduced NARU, a new benchmark designed to evaluate the capabilities of multimodal large language models (MLLMs) in understanding narrative evolution and cultural nuances within Japanese long-form videos. The benchmark comprises 1,481 questions based on 155 videos totaling 146.8 hours, covering four narrative and five cultural dimensions. Its construction involved a hierarchical memory-based annotation pipeline and verification by 68 native-speaking annotators. Initial evaluations indicate significant limitations in current MLLMs for long-range narrative integration and culturally grounded reasoning. AI
IMPACT NARU benchmark highlights critical gaps in MLLMs' ability to interpret complex narratives and cultural context in long-form video.
RANK_REASON The cluster describes a new academic benchmark for evaluating AI models. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →