PulseAugur
实时 09:46:31
English(EN) MELON: A Large-Scale Dataset for Multi-Event Text-to-Long-Video Retrieval

新数据集和模型推动文本到长视频检索能力发展

研究人员推出了MELON,这是一个专为多事件文本到长视频检索设计的新大规模数据集,解决了现有数据集主要关注具有单一事件的短片段的局限性。MELON标注了长视频中的多个事件间隔,使模型能够更好地理解复杂的多事件结构。同时,还提出了一个名为TAME的框架,通过将时间建模集成到基于CLIP的架构中来增强文本视频检索。TAME利用专家混合层和帧-时间标记来捕捉局部帧细节和长距离时间依赖性,在标准的检索基准上表现出改进的性能。 AI

影响 文本视频检索的进步可能导致更复杂的搜索和推荐系统。

排序理由 两篇研究论文介绍了用于文本视频检索的新数据集和框架。

在 arXiv cs.IR (Information Retrieval) 阅读 →

AI 生成摘要 · Google Gemini · 来自 3 个来源。 我们如何撰写摘要 →

新数据集和模型推动文本到长视频检索能力发展

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
两篇研究论文介绍了用于文本视频检索的新数据集和框架。
Source corroboration
3 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
paper, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
15 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+1 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准

报道来源 [3]

  1. Hugging Face Daily Papers TIER_1 English(EN) ·

    TAME:用于文本-视频检索的时间感知混合专家模型

    Text-Video Retrieval (TVR) retrieves videos that match a natural-language query, but extending image-text models such as CLIP to videos is fundamentally limited by the lack of temporal modeling. Videos exhibit frame-wise heterogeneity in appearance and motion, and compressing all…

  2. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · KyungTae Lim ·

    MELON:用于多事件文本到长视频检索的大规模数据集

    Existing text-video retrieval datasets primarily consist of short-form clips containing a single dominant event. While suitable for measuring basic vision-language alignment, they are limited in capturing real-world retrieval scenarios, where long-form videos naturally contain mu…

  3. arXiv cs.CV TIER_1 English(EN) · Uicheol Jung, Juyoung Hong, Hojung Kwon, Yukyung Choi ·

    TAME:用于文本-视频检索的时间感知混合专家模型

    arXiv:2609.02204v1 Announce Type: new Abstract: Text-Video Retrieval (TVR) retrieves videos that match a natural-language query, but extending image-text models such as CLIP to videos is fundamentally limited by the lack of temporal modeling. Videos exhibit frame-wise heterogenei…