MSVD-Turkish: a comprehensive multimodal video dataset for integrated vision and language research in Turkish
PulseAugur coverage of MSVD-Turkish: a comprehensive multimodal video dataset for integrated vision and language research in Turkish — every cluster mentioning MSVD-Turkish: a comprehensive multimodal video dataset for integrated vision and language research in Turkish across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
New WSV framework improves zero-shot video captioning with synthetic video generation
Researchers have developed a new framework called WSV for zero-shot video captioning that addresses the cross-modal gap between text-only training and video-based inference. The method involves generating synthetic vide…
-
New framework boosts video captioning accuracy without retraining models
Researchers have developed ProCap, a novel framework designed to enhance video captioning without retraining existing large vision-language models. This method uses a lightweight scoring mechanism to identify and priori…
-
New framework aligns text and video distributions for improved retrieval
Researchers have introduced the Distribution-Alignment Bridge (DAB), a novel framework for text-to-video retrieval that treats the task as a distribution alignment problem. Instead of direct matching, DAB models text an…