Researchers have developed PRISM-AH, a novel framework designed to recognize ambivalence and hesitancy in videos by analyzing multimodal data streams. This system models these conflicting affective states as a temporal conflict across facial, vocal, linguistic, and bodily cues. PRISM-AH aligns these modalities within short time windows and uses a streaming model to detect cross-modal dissonance and predict future states, incorporating participant metadata. A knowledge-guided large language model then reasons over structured evidence, with its output fused late in the process to improve performance. The framework achieved a macro F1 score of 0.6133 on a public test set, significantly outperforming a zero-shot baseline. AI
IMPACT This framework could improve the accuracy of affective state recognition in videos, with potential applications in mental health monitoring and human-computer interaction.
RANK_REASON This is a research paper detailing a new AI framework and its performance on a specific task. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX Code Finder for Papers
- Connected Papers
- CORE Recommender
- DagsHub
- Gotit.pub
- Hugging Face
- Influence Flower
- Litmaps
- Podakanti Satyajith Chary
- PRISM-AH
- ScienceCast
- scite Smart Citations
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →