Researchers have introduced Congruency Score (CS), a new method to better calibrate the congruence between long text descriptions and image content in vision-language models. This approach maps similarity evidence into a bounded score, addressing the modality gap between image and text embeddings. Evaluations on datasets like DOCCI and Urban1k showed that while post-hoc calibration maintains strong association with human judgments, projection-based methods can improve threshold calibration at the expense of retrieval performance. The study emphasizes treating long-text image-text congruence scoring as a distinct problem with multiple objectives: retrieval performance, human association, and threshold calibration. AI
IMPACT Improves the accuracy and interpretability of vision-language models for tasks involving detailed image-text matching.
RANK_REASON Research paper introducing a new method for vision-language models. [lever_c_demoted from research: ic=1 ai=1.0]
- Alessandro Gambetti
- alphaXiv
- arXiv
- CatalyzeX
- Congruency Score (CS)
- DagsHub
- DOCCI
- Gotit.pub
- Hugging Face
- ScienceCast
- Urban1k
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →