English(EN)Evidence-Grounded Multimodal Knowledge Graph Construction for Multi-Lecture Educational Reasoning
新基准和方法推动AI多模态推理发展
作者PulseAugur 编辑部·[4 个来源]·
研究人员正在开发用于多模态知识图谱补全和推理的新方法,将视觉语言模型(VLM)与图结构相结合。ViSR-KGC提出了一种视觉子图推理方法,结合了表示学习和VLM分析来推断缺失的实体。Q-CueGraph专注于查询条件下的视觉证据图谱,以指导VLM在多模态推理任务中检查图像。此外,还开发了一个从讲座视频构建证据支撑的多模态知识图谱的流程,并引入了一个名为GraphVerse的综合基准来评估多模态大语言模型在视觉图谱推理方面的能力。
AI
arXiv:2608.05833v1 Announce Type: new Abstract: Knowledge graph completion (KGC) aims to infer missing entities or relations from incomplete graph structures, and has evolved into multimodal knowledge graph completion (MMKGC), where entities are associated with multiple modalitie…
arXiv:2608.04452v1 Announce Type: cross Abstract: High-resolution pixels and crop or zoom tools give multimodal large language models the ability to inspect an image, but they do not provide a reliable task-conditioned policy for deciding where to inspect. Q-CueGraph makes this d…
arXiv:2608.03161v1 Announce Type: new Abstract: Lecture videos distribute knowledge across speech, slide text, diagrams, equations, and presentation order, which transcript-only retrieval does not fully preserve. This paper presents an evidence-grounded multimodal pipeline that t…
arXiv cs.CV
TIER_1English(EN)·Yuanfu Sun, Yuanhang Ren, Kang Li, Chuanhao Ji, Jiaxi Li, Jiajin Liu, Ninghao Liu, Qiaoyu Tan·
arXiv:2608.06769v1 Announce Type: new Abstract: Recent Multimodal Large Language Models (MLLMs) have achieved remarkable progress across diverse vision-language tasks, creating an urgent need for more challenging benchmarks. Yet existing evaluations still provide limited insight …