Researchers from Apple Machine Learning Research have identified significant vulnerabilities in Reinforcement Learning (RL)-finetuned Vision-Language Models (VLMs). While RL finetuning improves performance on visual reasoning benchmarks, these models exhibit a pronounced susceptibility to textual perturbations, such as misleading captions or incorrect chain-of-thought (CoT) traces. The study found that closed-source models, despite similar failure modes, maintain greater robustness and reasoning consistency compared to their open-source counterparts. This suggests a limitation in current open-source RL finetuning techniques rather than an inherent task constraint, highlighting an accuracy-faithfulness trade-off where benchmark accuracy gains can erode the reliability of CoT reasoning. AI
IMPACT Highlights limitations in current open-source RL finetuning for VLMs, suggesting a need for training and assessment protocols that emphasize faithfulness and robustness alongside accuracy.
RANK_REASON The cluster contains a research paper detailing findings on the robustness and consistency of RL-finetuned VLMs. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Apple Machine Learning Research →
- Anshul Shah
- Apple Machine Learning Research
- Arnab Mondal
- Harvard University
- Joerg Liebelt
- OpenAI
- Rosie Zhao
- Xiaoyu Zhu
- Xinke Deng
- Yang Yang
- Zhongyu Jiang
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →