PulseAugur
实时 12:51:10
English(EN) FiRE: Enhancing MLLMs with Fine-Grained Context Learning for Complex Image Retrieval

FiRE 通过细粒度上下文学习增强多模态大语言模型以进行复杂图像检索 · 跟踪 3 个来源

研究人员开发了 FiRE,这是一种增强多模态大语言模型(MLLMs)以进行复杂图像检索任务的新颖方法。FiRE 引入了一种细粒度上下文学习策略,该策略涉及一个两阶段的微调过程,区分推理和检索目标。该方法还包括一个自动化的管道,用于构建专门针对组合图像检索(CIR)的综合数据集。实验表明,即使使用资源消耗较少的 MLLM 主干,FiRE 在零样本检索设置下的表现也显著优于现有方法。 AI

影响 这项研究可能带来更复杂的图像搜索和多模态理解能力。

排序理由 该集群描述了一篇详细介绍增强 MLLMs 进行图像检索的新颖方法的最新研究论文。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 3 个来源。 我们如何撰写摘要 →

FiRE 通过细粒度上下文学习增强多模态大语言模型以进行复杂图像检索 · 跟踪 3 个来源

报道来源 [3]

  1. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Xiangyu Zhao ·

    FiRE:通过细粒度上下文学习增强多模态大模型以实现复杂图像检索

    Due to their strong generalizable multimodal processing and reasoning capabilities, Multimodal Large Language Models (MLLMs) have demonstrated significant potential as universal image retrievers, effectively addressing diverse real-world image retrieval tasks. Nevertheless, pione…

  2. Hugging Face Daily Papers TIER_1 English(EN) ·

    FiRE:通过细粒度上下文学习增强多模态大模型以实现复杂图像检索

    Due to their strong generalizable multimodal processing and reasoning capabilities, Multimodal Large Language Models (MLLMs) have demonstrated significant potential as universal image retrievers, effectively addressing diverse real-world image retrieval tasks. Nevertheless, pione…

  3. arXiv cs.CV TIER_1 English(EN) · Bohan Hou, Haoqiang Lin, Xuemeng Song, Haokun Wen, Meng Liu, Yupeng Hu, Xiangyu Zhao ·

    FiRE:通过细粒度上下文学习增强多模态大模型以进行复杂图像检索

    arXiv:2607.27959v1 Announce Type: new Abstract: Due to their strong generalizable multimodal processing and reasoning capabilities, Multimodal Large Language Models (MLLMs) have demonstrated significant potential as universal image retrievers, effectively addressing diverse real-…