Researchers have developed a new reinforcement learning strategy called Refusal-Calibrated Group Relative Policy Optimization (RC-GRPO) to improve the ability of Multimodal Large Language Models (MLLMs) to correctly identify when an object described in text does not exist. Current MLLMs often struggle with this, producing hallucinated outputs due to a lack of negative training samples. RC-GRPO aims to enhance refusal capabilities without sacrificing the model's core object localization accuracy on positive samples. Experiments on three benchmarks show that RC-GRPO achieves a better balance between accuracy and reliability. AI
IMPACT Enhances the reliability of MLLMs by improving their ability to handle negative cases, potentially reducing hallucinations in visual-grounding tasks.
RANK_REASON The cluster describes a new research paper detailing a novel method for improving MLLM capabilities. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- Generalized Referring Expression Comprehension
- MLLMs
- RC-GRPO
- Refusal-Calibrated Group Relative Policy Optimization
- reinforcement learning
- supervised fine-tuning
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →