Researchers have developed a new framework called Confidence-Aware On-Policy Distillation (CA-OPD) to improve autoregressive vision-language models. This method addresses compounding errors by using teacher confidence to correct unreliable student predictions during training. CA-OPD aligns knowledge transfer with intervention decisions, providing direct supervision at corrected positions and full predictive distribution at retained positions. When applied to GUI grounding and optical character recognition tasks, CA-OPD significantly enhanced the Qwen3.5-0.8B baseline, showing notable gains on benchmarks like ScreenSpot-Pro and OCRBench-v2 English. AI
IMPACT Enhances the performance of autoregressive vision-language models, potentially improving accuracy in tasks like GUI grounding and OCR.
RANK_REASON The cluster describes a new method and framework published in an academic paper on arXiv. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- CA-OPD
- DagsHub
- Hugging Face
- OCRBench-v2 English
- On-Policy Distillation
- Qwen3.5-0.8B
- ScreenSpot-Pro
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →