Researchers have introduced SlotNarrative, a novel interface designed to make Video Large Language Models (Video-LLMs) more token-efficient. This system organizes videos into persistent object narratives using compact object-state tokens, which represent an object's identity and its state (appearance, geometry, visibility, trajectory) across different segments. Unlike previous methods that compress frame-wise features, SlotNarrative first groups visual features into object-like slots and then associates recurring observations with clip-level entries through a parameter-free memory. This approach allocates a fixed 144 visual-token positions for a frozen Video-LLM, regardless of the number of sampled frames, and has demonstrated a favorable accuracy-token trade-off on various datasets. AI
IMPACT This new interface could significantly reduce computational costs for video understanding tasks, enabling broader application of Video-LLMs.
RANK_REASON The item is an academic paper detailing a new technical approach for video language models. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyX
- Connected Papers
- CORE Recommender
- DagsHub
- Gotit.pub
- Hugging Face
- Influence Flower
- Litmaps
- ScienceCast
- scite Smart Citations
- SlotNarrative
- Video LLMs
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →