Researchers have introduced DIM-Fashion, a new benchmark designed to improve multi-turn fashion image retrieval by addressing the limitations of existing methods. Current approaches often assume a uniform intent transition, failing to account for diverse user behaviors like rollbacks. DIM-Fashion, comprising 26,000 multi-turn sessions from various fashion retrieval datasets, aims to capture these complex interactions. Additionally, the team developed FashionAM, a multimodal large language model and vision-language pre-training framework that directly processes conversational queries against visual embeddings, bypassing intermediate textification to preserve finer visual details. AI
IMPACT Enhances multimodal understanding for complex, interactive search scenarios, potentially improving e-commerce and personalized recommendation systems.
RANK_REASON Research paper introducing a new benchmark and model. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- Connected Papers
- DagsHub
- DIM-Fashion
- FashionAM
- Gotit.pub
- Hugging Face
- Litmaps
- ScienceCast
- scite Smart Citations
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →