Researchers have developed VDC-Agent, a novel framework that enables a single multimodal large language model to autonomously generate and refine video detailed captions. This self-evolving system overcomes the reliance on costly human annotations or external models by employing a principle-guided self-reflection process. To enhance efficiency, the model's reflective capabilities are internalized through a Curriculum Direct Preference Optimization strategy, which uses a preference dataset derived from the agent's self-scored trajectories. Experiments show VDC-Agent achieves state-of-the-art performance on VDC and DREAM-1K benchmarks, improving caption detail and faithfulness while maintaining inference efficiency. AI
IMPACT This framework could reduce the cost and effort required for generating high-quality video captions, potentially impacting content creation and analysis tools.
RANK_REASON The cluster contains a research paper detailing a new method and framework for video captioning. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- Curriculum Direct Preference Optimization
- DREAM-1K
- multimodal large language model
- Qiang Wang
- VDC-Agent
- VDC-Agent-19K
- Vine Copula
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →