Researchers have developed VideoNorms, a new dataset designed to evaluate the cultural awareness of Video Large Language Models (VideoLLMs). The dataset includes over 3,000 human judgments derived from popular US and Chinese TV shows, focusing on the prediction of cultural norm adherence or violation and the identification of supporting evidence. Findings indicate that current VideoLLMs perform worse on Chinese cultural norms compared to US norms, and struggle more with identifying non-verbal evidence. The study also suggests that while video modality is crucial, simply scaling up model size does not necessarily improve performance on this task. AI
IMPACT Highlights the need for culturally-aware AI training and evaluation, potentially guiding future development of more globally applicable VideoLLMs.
RANK_REASON The cluster is about an academic paper introducing a new dataset and benchmark for evaluating AI models. [lever_c_demoted from research: ic=1 ai=1.0]
- Arkadiy Saakyan
- arXiv
- Category:Television shows based on South Korean webtoons
- Hugging Face
- Standard Chinese
- US
- VideoLLMs
- VideoNorms
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →