PulseAugur
中
实时 16:50:04
English(EN) TiTok: Audio-Visual LLM for Multi-Segment Temporal Grounding

TiTok 视听大语言模型可精确地定位多个时间事件片段

研究人员推出 TiTok,这是一种视听大语言模型,旨在精确识别未修剪视频中的多个时间事件片段。该模型采用新颖的时间令牌交织(TTI)方法,通过将特殊时间令牌集成到视听流中来增强边界预测。为解决计数校准不准的问题,TiTok 采用了解耦奖励系统,并使用组奖励-解耦归一化策略优化(GDPO)进行优化,在新评估协议上取得了 65.7 mIoU 和 0.58 CountF1 的最先进成果。 AI

影响 引入了一种新颖的视听时间定位方法,有望改进视频分析和内容检索系统。

排序理由 详细介绍新模型和方法的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CV 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

TiTok 视听大语言模型可精确地定位多个时间事件片段

本文如何被排名

Signal score
4 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
详细介绍新模型和方法的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Same-day
Cluster formed today. Ranking reflects the current source set at time of score.

完整方法见我们的编辑标准。

报道来源 [1]

  1. arXiv cs.CV TIER_1 English(EN) · Eunji Shin, Dahyun Choi, Seungyeon Jo, Yejin Hong, Jiyoung Lee ·

    TiTok: 用于多段时间定位的视听大语言模型

    arXiv:2610.09408v1 Announce Type: new Abstract: Audio-visual multi-segment grounding (AV-MSG) in untrimmed videos, reasoning over audio-visual evidence and predicting multiple segments for a query, is a fundamental problem but remains challenging. Visual-only models overlook comp…