Two new research papers propose novel methods for compressing video tokens in Video Large Language Models (Video-LLMs) to improve efficiency. The first paper introduces NovaCov, a set-wise token compressor designed for streaming video that uses a historical reference bank and a dual-branch submodular coverage objective to retain representative content while reducing LLM prefilling latency by 46%. The second paper presents ONCE, a framework that learns a frequency-aware global codebook offline and reuses it for lightweight online compression, significantly reducing per-video computation and inference latency. AI
IMPACT These methods aim to reduce inference latency and computational costs for Video-LLMs, potentially enabling wider adoption and more efficient processing of video content.
RANK_REASON Two academic papers published on arXiv introducing novel methods for video token compression in Video-LLMs.
- alphaXiv
- arXiv
- CatalyzeX
- Connected Papers
- DagsHub
- Gotit.pub
- Hugging Face
- Influence Flower
- Litmaps
- NovaCov
- ONCE
- ScienceCast
- scite Smart Citations
- Video-LLMs
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →