A new research paper explores how action post-training affects the depth perception capabilities of vision-language models (VLMs). The study found that this post-training process, used to build vision-language-action (VLA) models, significantly degrades depth decodability across all layers compared to the base VLM. This degradation is particularly pronounced in the late layers, a phenomenon termed the 'cliff,' which is causally linked to interference within the late-layer MLPs. AI
IMPACT This research highlights potential trade-offs in VLM training, suggesting that action-oriented fine-tuning may compromise core spatial understanding capabilities.
RANK_REASON The cluster contains a research paper detailing findings on VLM capabilities. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- Gotit.pub
- Hugging Face
- Molmo2-ER
- MolmoAct2-LIBERO
- multilayer perceptron
- ScienceCast
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →