PulseAugur
实时 09:30:58
English(EN) CrossModalQA: A Cross-modal and Multi-hop Benchmark for Multimodal Retrieval-augmented Generation

新的CrossModalQA基准测试用于评估多模态大语言模型在复杂推理方面的能力

研究人员推出CrossModalQA,一个旨在评估多模态大语言模型(MLLMs)在检索增强生成(RAG)任务中的新基准测试。该基准测试通过专注于开放域证据发现以及跨文本和图像的复杂多跳推理,解决了现有系统中的局限性。CrossModalQA包含超过1,800个问答对,源自Wikipedia和Wikimedia Commons,要求模型执行平均跳数达3.50的多跳、跨模态推理。初步实验表明,当前MLLMs在完整证据检索方面存在困难,并且性能受到不完整或干扰性上下文的显著影响。 AI

影响 该基准测试将推动更强大的多模态人工智能系统的发展,使其能够进行复杂的推理和证据检索。

排序理由 该集群描述了一个用于评估AI模型的新学术基准测试。

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的CrossModalQA基准测试用于评估多模态大语言模型在复杂推理方面的能力

本文如何被排名

Signal score
13 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群描述了一个用于评估AI模型的新学术基准测试。
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Jiacheng Cai, Zijin Hong, Zheng Yuan, Huachi Zhou, Qinggang Zhang, Xiao Huang ·

    CrossModalQA:一个用于多模态检索增强生成的跨模态和多跳基准

    arXiv:2609.05518v1 Announce Type: cross Abstract: Despite the strong capabilities of multimodal large language models (MLLMs), their parametric knowledge remains incomplete and difficult to update, motivating multimodal retrieval-augmented generation (RAG) to ground responses in …