Researchers have developed a new framework called WSV for zero-shot video captioning that addresses the cross-modal gap between text-only training and video-based inference. The method involves generating synthetic video latent representations using a text-to-video model, which are then refined by a polisher to improve fidelity. Finally, a prompter conditions GPT-2 on these polished representations to generate captions. This approach achieved scores of 52 on B@4 and 95.7 on CIDEr metrics across MSVD, MSR-VTT, and VATEX datasets. AI
IMPACT This research introduces a novel approach to bridge the gap in zero-shot video captioning, potentially improving the accuracy of generated captions by leveraging synthetic video data.
RANK_REASON The cluster contains a research paper detailing a novel framework for zero-shot video captioning. [lever_c_demoted from research: ic=1 ai=1.0]
- 3D Causal VAE
- arXiv
- GPT-2
- Hugging Face
- MSR-VTT
- MSVD-Turkish: a comprehensive multimodal video dataset for integrated vision and language research in Turkish
- VATEX
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →