A recent study examined the confidence signals of two small vision-language models, Qwen2-VL-2B-Instruct and SmolVLM-Instruct, under realistic image degradation. The research found a significant discrepancy between the models' stated confidence in natural language and their internal token probabilities. While internal probabilities proved to be a more reliable indicator of errors, both models struggled to accurately assess their confidence when faced with severe underexposure, leading to a collapse in accuracy with minimal change in confidence signals. AI
IMPACT Highlights limitations in current small vision-language models regarding reliable uncertainty estimation, crucial for safe deployment.
RANK_REASON Academic paper detailing research findings on model capabilities. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →