PulseAugur
EN
LIVE 08:25:58

VDC-Agent framework enables autonomous self-evolving video captioning

Researchers have developed VDC-Agent, a novel framework that enables a single multimodal large language model to autonomously generate and refine video detailed captions. This self-evolving system overcomes the reliance on costly human annotations or external models by employing a principle-guided self-reflection process. To enhance efficiency, the model's reflective capabilities are internalized through a Curriculum Direct Preference Optimization strategy, which uses a preference dataset derived from the agent's self-scored trajectories. Experiments show VDC-Agent achieves state-of-the-art performance on VDC and DREAM-1K benchmarks, improving caption detail and faithfulness while maintaining inference efficiency. AI

IMPACT This framework could reduce the cost and effort required for generating high-quality video captions, potentially impacting content creation and analysis tools.

RANK_REASON The cluster contains a research paper detailing a new method and framework for video captioning. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

VDC-Agent framework enables autonomous self-evolving video captioning

COVERAGE [1]

  1. arXiv cs.AI TIER_1 English(EN) · Qiang Wang, Xinyuan Gao, Yuhang He, Jizhou Han, Jiangyang Li, SongLin Dong, Zhiheng Ma, Yihong Gong ·

    VDC-Agent: When Video Detailed Captioners Evolve Themselves via Agentic Self-Reflection

    arXiv:2511.19436v2 Announce Type: replace-cross Abstract: Existing Video Detailed Captioning (VDC) methods predominantly rely on costly human annotations or distillation from powerful proprietary models, creating a dependency on external supervision. In this paper, we propose VDC…