PulseAugur
实时 10:14:02
English(EN) ShotFinder: Imagination-Driven Open-Domain Video Shot Retrieval via Web Search

新的ShotFinder基准揭示多模态大语言模型难以进行视频编辑

研究人员推出了ShotFinder,这是一个旨在评估大型语言模型开放域视频片段检索能力的新基准。该基准将编辑需求形式化为面向关键帧的片段描述,并包含五种可控约束:时间顺序、颜色、视觉风格、音频和分辨率。从YouTube收集了1,210个样本的数据集,并提出了一个三阶段检索和定位流程。实验表明,当前模型与人类能力之间存在显著的性能差距,其中颜色和视觉风格对多模态大型模型构成了最大的挑战。 AI

影响 凸显了多模态大语言模型在复杂视频编辑任务中的当前局限性,指出了未来研究和发展的方向。

排序理由 该项目描述了一个新的基准和相关论文,用于评估AI在视频片段检索方面的能力。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的ShotFinder基准揭示多模态大语言模型难以进行视频编辑

本文如何被排名

Signal score
12 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该项目描述了一个新的基准和相关论文,用于评估AI在视频片段检索方面的能力。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Same-day
Cluster formed today. Ranking reflects the current source set at time of score.

完整方法见我们的编辑标准

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Tao Yu, Haopeng Jin, Hao Wang, Shenghua Chai, Yujia Yang, Junhao Gong, Jiaming Guo, Minghui Zhang, Xinlong Chen, Zhenghao Zhang, Yuxuan Zhou, Yufei Xiong, Shanbin Zhang, Jiabing Yang, YiFan Zhang, Hongzhu Yi, Xinming Wang, Cheng Zhong, Xiao Ma, Zhang Zha… ·

    ShotFinder:通过网络搜索实现由想象驱动的开放域视频镜头检索

    arXiv:2601.23232v4 Announce Type: replace-cross Abstract: In recent years, large language models (LLMs) have made rapid progress in information retrieval, yet existing research has mainly focused on text or static multimodal settings. Open-domain video shot retrieval, which invol…