Researchers have introduced NARU, a new benchmark designed to evaluate the capabilities of multimodal large language models (MLLMs) in understanding narrative evolution and cultural nuances within Japanese long-form videos. The benchmark comprises 1,481 questions derived from 155 videos, totaling over 146 hours, and assesses four narrative and five cultural dimensions. Developed through a hierarchical annotation pipeline involving native speakers, NARU aims to identify limitations in current MLLMs' ability to process complex, high-context video content. AI
IMPACT This benchmark aims to improve AI's ability to understand complex, culturally nuanced long-form video content.
RANK_REASON The cluster describes a new academic benchmark for evaluating AI models on a specific type of data (Japanese long-form video). [lever_c_demoted from research: ic=1 ai=1.0]
Read on Hugging Face Daily Papers →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →