ActivityNet Captions
PulseAugur coverage of ActivityNet Captions — every cluster mentioning ActivityNet Captions across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
New framework improves video captioning using VLM-guided transition discovery
Researchers have developed a new framework called Seeing Before Synthesizing (SBS) to improve weakly-supervised dense video captioning. This method uses vision-language models (VLMs) to generate frame-level narratives f…
-
New method improves VLM temporal grounding by asking binary questions
Researchers have developed a novel training-free method called FV-Action for temporal grounding in vision-language models (VLMs). This approach addresses the issue of VLMs confidently providing incorrect timestamps for …
-
AI framework enhances semantic video communication with new routing
Researchers have developed a generative AI framework for semantic video communication, aiming to transmit meaning rather than raw data. The system addresses challenges in temporal modeling for bandwidth constraints and …
-
GenSpan framework improves video retrieval for complex action queries
Researchers have developed GenSpan, a new framework for video corpus moment retrieval that specifically addresses challenges with multi-verb queries. GenSpan utilizes auxiliary videos generated from subtitle cues to act…
-
New network transfers knowledge for unsupervised video-text matching
Researchers have developed a novel cross-modal knowledge transfer network for unsupervised temporal sentence grounding. This approach aims to overcome the reliance on expensive, paired video-query annotations by leverag…
-
PEEK method efficiently selects key video frames for captioning
Researchers have developed PEEK, an efficient method for selecting essential frames from videos for captioning. This technique distills knowledge from a larger teacher model into a smaller one, enabling it to identify t…