PulseAugur
EN
LIVE 10:49:15

New NARU benchmark tests MLLMs on Japanese video narrative and culture

Researchers have introduced NARU, a new benchmark designed to evaluate the capabilities of multimodal large language models (MLLMs) in understanding narrative evolution and cultural nuances within Japanese long-form videos. The benchmark comprises 1,481 questions based on 155 videos totaling 146.8 hours, covering four narrative and five cultural dimensions. Its construction involved a hierarchical memory-based annotation pipeline and verification by 68 native-speaking annotators. Initial evaluations indicate significant limitations in current MLLMs for long-range narrative integration and culturally grounded reasoning. AI

IMPACT NARU benchmark highlights critical gaps in MLLMs' ability to interpret complex narratives and cultural context in long-form video.

RANK_REASON The cluster describes a new academic benchmark for evaluating AI models. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New NARU benchmark tests MLLMs on Japanese video narrative and culture

COVERAGE [1]

  1. arXiv cs.AI TIER_1 English(EN) · Yuheng Huang, Jianlang Chen, Jiayang Song, Hua Qi, Aza Kai, Vincent Markert, Edison Marrese-Taylor, Jianjun Zhao, Lei Ma ·

    NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video

    arXiv:2608.13210v1 Announce Type: cross Abstract: Long-form video understanding encompasses tasks that go beyond retrieving isolated events, including tracking an evolving narrative and interpreting social meaning that may remain implicit. However, existing benchmarks rarely eval…