Researchers have developed MAG, a novel framework designed to enhance in-context learning for multi-modal large language models (MLLMs) by effectively utilizing unlabeled data. MAG addresses the challenge of selecting high-quality demonstrations, which significantly impacts MLLM performance, by treating demonstration selection as a semi-supervised problem on a multi-modal graph. The framework employs a two-stage process: first, it propagates relevance scores through unlabeled data to identify impactful samples for pseudo-labeling, thereby reducing computational costs. Second, it uses multi-modal relevance, incorporating both visual and textual information, to finalize the demonstration selection. Experiments across eight multi-modal benchmarks show that MAG consistently surpasses existing methods in low-label scenarios, demonstrating substantial improvements with a constrained pseudo-labeling budget. AI
IMPACT Enhances multi-modal LLM adaptability in low-data scenarios, potentially improving performance across various cross-modal tasks.
RANK_REASON The cluster contains a research paper detailing a new framework for multi-modal in-context learning. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- eight multi-modal benchmarks
- Hugging Face
- MAnifold Guided Semi-Supervised Multi-modal In-Context Learning
- MLLMs
- Multi-modal graph regularization based class center discriminant analysis for cross modal retrieval
- Multi-modal Large Language Models
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →