Researchers have introduced CoCo-IR, a novel task and model for contextual composed image retrieval that allows for iterative refinement of visual searches. The proposed Large Multimodal Model (LMM) interprets interaction history to generate evolving Transformable Image Embeddings (TIE). To facilitate training, an autonomous data engine leverages LMMs for data generation and model-guided verification for hard negatives. CoCo-IR demonstrates state-of-the-art performance, achieving 39.4 mAP@5 on the CIRCO benchmark and 44.1 R@1 on its new 4-turn dialogue benchmark, significantly outperforming existing methods. AI
IMPACT Enables more sophisticated, multi-turn visual search capabilities, potentially improving user experience in image-based applications.
RANK_REASON The cluster describes a new research paper introducing a novel task and model for image retrieval.
Read on Hugging Face Daily Papers →
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →