Researchers have investigated how the structure of controlled vocabularies, like the Art and Architecture Thesaurus (AAT), impacts the performance of vision-language models such as CLIP when retrieving images from historical photo collections. The study found that properties like root facet type and hierarchy depth significantly influence CLIP's success and failure modes, with visual coherence and text-image alignment being key differentiators. While standard retrieval metrics did not strongly correlate with these structural properties, fine-tuning the model showed improvements, particularly for shallower terms in the hierarchy. AI
IMPACT Provides insights into improving image retrieval for historical archives using AI models.
RANK_REASON Academic paper detailing research findings on AI model performance. [lever_c_demoted from research: ic=1 ai=1.0]
Read on arXiv cs.IR (Information Retrieval) →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →