Two new research papers propose methods for compressing video data for multimodal large language models (MLLMs). The first paper, Visual Token Coding (VTC), uses principles from classical video coding to predict frames and measure redundancy, achieving 100.1% performance retention with a 50% token budget on Qwen3-VL. The second paper, Token-Budget Distillation (TBD), employs a dual-path teacher-student design to transfer full-token semantics to compressed video VLMs, maintaining 97.0% accuracy on LLaVA-Video with a 10% token budget. Both methods aim to reduce the computational cost of processing video inputs for MLLMs without significant performance degradation. AI
IMPACT These techniques could significantly reduce the computational cost and improve the efficiency of video understanding in multimodal AI systems.
RANK_REASON Two research papers published on arXiv proposing new methods for video token compression in multimodal LLMs.
- arXiv
- LLaVA-onevision
- LLaVA-Video
- Qwen3-VL
- Qwen3-VL-8B-Instruct
- Token-Budget Distillation
- Visual Token Coding
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →