Researchers have developed a new framework called the Capability-Driven Multimodal Scaling Law to predict the performance of vision-language models (VLMs) before training. This law uses a low-dimensional capability score extracted from LLM textual benchmarks to model VLM accuracy, considering transfer and absorption rates specific to each backbone model. Experiments training over 150 VLMs with 34 LLMs demonstrated that this framework can accurately extrapolate performance from smaller models to larger ones and predict entire training trajectories. The research also revealed that certain textual benchmarks can negatively correlate with multimodal performance, indicating benchmark-gaming, and that base LLMs are often better VLM backbones than instruction-tuned models. AI
IMPACT Provides a principled, quantitative method for selecting LLM backbones for VLMs, potentially saving significant compute and time.
RANK_REASON The cluster contains an academic paper detailing a new framework and experimental results for predicting VLM performance. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- Capability-Driven Multimodal Scaling Law
- CatalyzeX
- DagsHub
- Gotit.pub
- Hugging Face
- LLM
- principal component analysis
- ScienceCast
- Vision--Language Models
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →