Researchers have introduced Behavior Pack Optimization (BPO), a novel post-training method for video multimodal large language models (MLLMs). BPO addresses the issue of MLLMs relying on appearance and language priors rather than temporal evidence by computing rewards across a pack of outputs from counterfactual views. This approach encourages models to be stable when irrelevant interventions occur, sensitive when key evidence is removed, and to abstain when no evidence is present. Applied to Qwen2.5-VL-7B-Instruct, BPO significantly improved accuracy on benchmarks like TempCompass, MVBench, and NExT-QA, with gains also observed in Video-MME, LongVideoBench, and LLaVA-Video-7B. AI
IMPACT Enhances video LLM reasoning by improving temporal evidence utilization, potentially leading to more reliable video understanding.
RANK_REASON Academic paper detailing a new method for improving video multimodal large language models. [lever_c_demoted from research: ic=1 ai=1.0]
- Behavior Pack Optimization
- GRPO
- LLaVA-Video-7B
- LongVideoBench
- MVBench
- NExT-QA
- Qwen2.5-VL-7B-Instruct
- TempCompass
- Video-MME
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →