Researchers have introduced TiTok, an audio-visual large language model designed to precisely identify multiple temporal event segments within untrimmed videos. The model employs a novel Time Token Interleaving (TTI) method to enhance boundary prediction by integrating special time tokens into the audio-visual stream. To address count miscalibration, TiTok utilizes a decoupled reward system optimized with Group reward-Decoupled Normalization Policy Optimization (GDPO), achieving state-of-the-art results with 65.7 mIoU and 0.58 CountF1 on a new evaluation protocol. AI
IMPACT Introduces a novel approach to audio-visual temporal grounding, potentially improving video analysis and content retrieval systems.
RANK_REASON Academic paper detailing a new model and methodology. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- CountF1
- Group reward-Decoupled Normalization Policy Optimization
- Hugging Face
- Time Token Interleaving
- UnAV-100
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →