A new research paper published on arXiv questions the reliability of split-conformal prediction as a safety measure for zero-shot vision-language models (VLMs) when deployed under shifting data conditions. The study found that while marginal coverage can remain high, class-conditional tail coverage can significantly degrade, with some classes experiencing near-zero coverage. Various calibration techniques were tested, with target-side class calibration showing the most promise but requiring extensive labeled data. AI
IMPACT Highlights potential safety risks in current VLM deployment strategies, suggesting a need for more robust calibration methods.
RANK_REASON Research paper published on arXiv discussing model safety and calibration techniques. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- clustered conformal
- Conf-OT
- ImageNet
- ImageNet-Sketch
- Mondrian calibration
- SigLIP
- zero-shot vision-language models
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →