A new survey paper systematically categorizes the field of Efficient Multimodal Learning (EML), addressing computational and memory bottlenecks in multimodal models. It proposes a model-to-system taxonomy, analyzing over 300 works across three hierarchical levels: model, algorithm, and system. The paper synthesizes insights on the trade-offs between efficiency, utility, and privacy, using Multimodal Large Language Models (MLLMs) as a case study to illustrate the field's evolution and future directions towards intrinsically efficient AI. AI
影响 Provides a structured framework for understanding and developing efficient multimodal AI systems, potentially accelerating deployment.
排序理由 The item is a survey paper published on arXiv detailing research in Efficient Multimodal Learning. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX Code Finder for Papers
- Connected Papers
- CORE Recommender
- DagsHub
- Efficient Multimodal Learning
- Gotit.pub
- Hugging Face
- Influence Flower
- Litmaps
- Multimodal Large Language Models and Tunings: Vision, Language, Sensors, Audio, and Beyond
- ScienceCast
- scite Smart Citations
AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →