A new survey paper published on arXiv details various inference-efficiency mechanisms for Video Large Language Models (VideoLLMs). These models, which combine video representations with large language models, face significant computational and memory costs that limit their deployment in resource-constrained environments. The survey categorizes methods by pipeline stage, including frame sampling, modality encoding, and LLM prefilling, and analyzes their impact on parameter count, FLOPs, latency, and memory usage. It also highlights gaps in audiovisual efficiency and the need for standardized evaluation, while maintaining a repository of relevant research. AI
IMPACT Identifies key areas for optimizing the computational cost of video-based AI models, potentially enabling wider deployment.
RANK_REASON The item is a survey paper published on arXiv detailing technical mechanisms for improving AI model efficiency. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- Connected Papers
- CORE Recommender
- DagsHub
- Gotit.pub
- Hugging Face
- Influence Flower
- Litmaps
- ScienceCast
- scite Smart Citations
- VideoLLMs
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →