Researchers have developed GAD-RL, a novel method to improve the faithfulness of optical character recognition (OCR) in vision-language models. This technique adaptively regulates teacher supervision during post-training, adjusting based on the student model's performance and response distributions. GAD-RL aims to prevent models from rewriting anomalous text into plausible but incorrect expressions, thereby enhancing transcription accuracy. The method showed significant improvements, achieving 59.92% Micro Recall on CHAOS-Bench when applied to the Qwen3.5-2B model. AI
IMPACT Enhances the accuracy of OCR in vision-language models, crucial for applications relying on accurate text extraction from images.
RANK_REASON The cluster contains an academic paper detailing a new method and benchmark results. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →