Researchers have developed a new framework called Caption-once, Frames-on-Demand (CFD) designed for efficient long-video understanding on edge devices. This system utilizes a dual-track narrative index, combining an event-level story skeleton with a clip-level micro-log, to reduce the need for re-captioning. A cloud-side multimodal large language model (MLLM) then uses a Visual-Need Router to selectively retrieve keyframes for perceptual queries, while keeping temporal-structural questions within the language domain, thereby optimizing compute and bandwidth usage. AI
IMPACT This approach could significantly improve the feasibility of analyzing long video content on resource-constrained devices.
RANK_REASON The item describes a novel framework and methodology presented in an academic paper. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- Caption-once, Frames-on-Demand: Visual-Need Routing for Budget-Aware Agentic Long Video Understanding
- Hugging Face
- multimodal large language model
- Visual-Need Router
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →