Researchers have developed Video2Reaction, a new dataset designed to train video foundation models to predict audience emotional responses. This dataset maps short video clips to viewer reactions expressed through social media comments, modeling emotions as distributions to capture subjectivity. Benchmarking finetuned vision-language models (VLMs) like LLaVA-NeXT-Video-7B, the study shows that VLMs trained on Video2Reaction can effectively predict dominant reactions and transfer this capability to other emotion-related datasets, achieving performance comparable to models trained on full datasets. AI
IMPACT This dataset could enable AI systems to better understand and respond to human emotional cues in video content.
RANK_REASON The cluster describes a new academic paper introducing a dataset and benchmarking models. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →