YouCook2
PulseAugur coverage of YouCook2 — every cluster mentioning YouCook2 across labs, papers, and developer communities, ranked by signal.
-
New framework improves video captioning using VLM-guided transition discovery
Researchers have developed a new framework called Seeing Before Synthesizing (SBS) to improve weakly-supervised dense video captioning. This method uses vision-language models (VLMs) to generate frame-level narratives f…
-
ClipSum framework uses CLIP for better instructional video summaries
Researchers have developed ClipSum, a new framework for summarizing instructional videos by leveraging CLIP's vision-language features. This approach uses semantically aligned visual features from CLIP, trained on a vas…
-
Grounding Video Reasoning in Physical Signals
Researchers have developed a new benchmark for evaluating physical video understanding, moving beyond simple event recognition to assess a model's ability to pinpoint events in time and space. This benchmark, which incl…