PulseAugur
实时 12:53:40

新方法以有限数据增强无监督跨模态检索 · 跟踪4个来源

研究人员正在开发新的无监督跨模态检索方法,旨在提高效率并减少对大型手动标注数据集的依赖。论文提出了属性提示核哈希(APKH)和全局邻域对齐哈希(GNAH)等技术,这些技术利用视觉语言基础模型和有限的配对数据来构建紧凑、对齐的汉明空间。另一种方法UniCA引入了双向交叉注意力和正相似性损失,以实现更鲁棒的多模态检索,并在WebQA+等基准测试中取得了改进。 AI

影响 这些研究工作旨在通过减少数据需求和改进对齐技术,使跨模态检索更加易于访问和高效。

排序理由 多篇学术论文提出新的跨模态检索方法。

在 arXiv cs.IR (Information Retrieval) 阅读 →

AI 生成摘要 · Google Gemini · 来自 6 个来源。 我们如何撰写摘要 →

新方法以有限数据增强无监督跨模态检索 · 跟踪4个来源

报道来源 [6]

  1. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Yap-Peng Tan ·

    用于无监督数据高效跨模态检索的属性提示核哈希

    Unsupervised cross-modal hashing enables efficient retrieval of semantically related instances across different modalities without requiring manual semantic annotation. However, existing unsupervised methods rely heavily on large-scale image-text pairs. Collecting such data can b…

  2. arXiv cs.AI TIER_1 English(EN) · Fan Xu, Luis A. Leiva ·

    跨模态信息检索的多模态表示对齐

    arXiv:2506.08774v2 Announce Type: replace-cross Abstract: Different machine learning models can represent the same underlying concept in different ways. This variability is particularly valuable for in-the-wild multimodal retrieval, where the objective is to identify the correspo…

  3. arXiv cs.CV TIER_1 English(EN) · Runhao Li, Xiaoxu Ma, Zhenyu Weng, Yue Zhang, Guibo Luo, Huiping Zhuang, Zhiping Lin, Yap-Peng Tan ·

    用于无监督数据高效跨模态检索的属性提示核哈希

    arXiv:2607.00379v1 Announce Type: cross Abstract: Unsupervised cross-modal hashing enables efficient retrieval of semantically related instances across different modalities without requiring manual semantic annotation. However, existing unsupervised methods rely heavily on large-…

  4. arXiv cs.CV TIER_1 English(EN) · Zixu Zhao, Yang Zhan, Yunhao Li, Yan Li ·

    TCMA:用于无人机跨模态文本-视频检索的文本条件化多粒度对齐

    arXiv:2510.10180v2 Announce Type: replace Abstract: Unmanned aerial vehicles (UAVs) have become powerful platforms for real-time, high-resolution data collection, producing massive volumes of aerial videos. Efficient retrieval of relevant content from these videos is crucial for …

  5. arXiv cs.CV TIER_1 English(EN) · Yap-Peng Tan ·

    无监督数据高效跨模态检索与全局邻域对齐哈希

    Compared to supervised cross-modal hashing (CMH), unsupervised CMH reduces the reliance on manual labeling by learning binary codes from unlabeled image-text pairs. However, existing unsupervised CMH methods often rely on large-scale image-text pairs, which are costly to collect.…

  6. arXiv cs.CV TIER_1 English(EN) · Yini Huang, Wenlong Zhang ·

    UniCA:具有正相似性损失的双向交叉注意力,用于鲁棒的多模态检索

    arXiv:2606.28350v1 Announce Type: cross Abstract: Multi-modal retrieval has become increasingly critical for handling the growing volume of integrated visual-textual data in real-world applications, but existing frameworks rely on implicit fusion via text encoder self-attention, …