PulseAugur
EN
LIVE 06:33:12

New dataset VideoNorms tests cultural awareness in VideoLLMs

Researchers have developed VideoNorms, a new dataset designed to evaluate the cultural awareness of Video Large Language Models (VideoLLMs). The dataset includes over 3,000 human judgments derived from popular US and Chinese TV shows, focusing on the prediction of cultural norm adherence or violation and the identification of supporting evidence. Findings indicate that current VideoLLMs perform worse on Chinese cultural norms compared to US norms, and struggle more with identifying non-verbal evidence. The study also suggests that while video modality is crucial, simply scaling up model size does not necessarily improve performance on this task. AI

IMPACT Highlights the need for culturally-aware AI training and evaluation, potentially guiding future development of more globally applicable VideoLLMs.

RANK_REASON The cluster is about an academic paper introducing a new dataset and benchmark for evaluating AI models. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.CL →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New dataset VideoNorms tests cultural awareness in VideoLLMs

COVERAGE [1]

  1. arXiv cs.CL TIER_1 English(EN) · Nikhil Reddy Varimalla, Yunfei Xu, Meng Fan Wang, Arkadiy Saakyan, Smaranda Muresan ·

    VideoNorms: Benchmarking Cultural Awareness of Video Language Models

    arXiv:2510.08543v2 Announce Type: replace-cross Abstract: As Video Large Language Models (VideoLLMs) are deployed globally, it is important to assess their ability to reason across cultural contexts. To advance cultural norm awareness evaluation in VideoLLMs, we introduce VideoNo…