Researchers have introduced M-Drama, a new benchmark designed to improve the understanding of micro-dramas, which are characterized by their extremely short duration and dense storylines. This benchmark includes over 35,000 instances across 9,138 clips and is bilingual. To enhance the performance of vision-language models (VLMs) on complex narratives, a novel graph-matching reward function called SAGA (Structure-Aware Graph Alignment) has been developed. SAGA models narratives as heterogeneous graphs and provides dense rewards through semantic triplet and structural temporal matching, outperforming existing methods on the Qwen3-VL-8B-Instruct model. AI
IMPACT This research could lead to more sophisticated AI models capable of understanding nuanced narratives in short-form video content.
RANK_REASON The item is an academic paper detailing a new benchmark and a novel method for a specific AI task. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX Code Finder for Papers
- Connected Papers
- CORE Recommender
- DagsHub
- Gotit.pub
- Hugging Face
- Influence Flower
- Litmaps
- M-Drama
- Qwen3-VL-8B-Instruct
- SAGA
- ScienceCast
- scite Smart Citations
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →