Researchers have introduced Chehre, a new dataset designed to explore perceptual variations in how video language models interpret facial expressions. The dataset comprises over 2,000 videos, with each video annotated by approximately 30 individuals, capturing a wide range of dynamic facial expressions prompted by emojis. This resource enables a new task called "distributional expression recognition," which evaluates a model's ability to replicate the diversity of human perception. Initial tests indicate that persona prompting can effectively guide model perception and better align it with human annotator responses. AI
IMPACT This dataset could lead to more robust video language models capable of understanding nuanced human expressions.
RANK_REASON The cluster describes a new dataset and research paper published on arXiv, focusing on a novel task for evaluating video language models. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →