PulseAugur
实时 19:55:54
English(EN) Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval

新的UniME-R1框架通过反馈驱动的推理改进多模态检索 · 跟踪2个来源

研究人员开发了UniME-R1,这是一个新颖的框架,旨在通过将检索反馈纳入推理过程来增强统一多模态检索。与以往仅依赖基于查询的思维链(CoT)的方法不同,UniME-R1的顾问分析初步检索到的候选对象,以识别混淆点并生成检索中心化思维链(RC-CoT)。这种方法可以优化检索方向,并在MMEB-V2等基准测试中提高性能。 AI

影响 通过反馈驱动的推理实现更准确的候选对象识别,从而增强多模态检索系统。

排序理由 该集群描述了一篇详细介绍多模态检索新颖框架的研究论文。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

新的UniME-R1框架通过反馈驱动的推理改进多模态检索 · 跟踪2个来源

报道来源 [2]

  1. Hugging Face Daily Papers TIER_1 English(EN) ·

    从失败中学习:通过硬负例进行以检索为中心的CoT,实现统一多模态检索

    Unified multimodal retrieval aims to identify candidates that satisfy complex user intent expressed through heterogeneous inputs. Although Large Vision-Language Model (LVLM)-based retrievers are efficient and scalable, directly encoding raw multimodal inputs often misses fine-gra…

  2. arXiv cs.CV TIER_1 English(EN) · Zelong Sun, Jun Wang, Kaicheng Yang, Tiancheng Gu, Ziyong Feng, Zhiwu Lu ·

    从失败中学习:通过硬负例进行以检索为中心的CoT以实现统一的多模态检索

    arXiv:2608.06060v1 Announce Type: new Abstract: Unified multimodal retrieval aims to identify candidates that satisfy complex user intent expressed through heterogeneous inputs. Although Large Vision-Language Model (LVLM)-based retrievers are efficient and scalable, directly enco…