KB-VQA
PulseAugur coverage of KB-VQA — every cluster mentioning KB-VQA across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
New methods enhance multimodal models for visual question answering · 6 sources tracked
Researchers have developed several new methods to improve the performance of multimodal large language models (MLLMs) in knowledge-based visual question answering (KB-VQA). One approach, 'Look Twice,' is a training-free…
-
MMAgent-R^2 enhances multi-modal retrieval with visual reranking and rejection · 2 sources tracked
Researchers have introduced MMAgent-R$^2$, a novel agentic framework designed to enhance multi-modal retrieval augmented generation (mRAG) systems. This framework addresses limitations in existing mRAG methods that stru…
-
ProMSA agent advances knowledge-based visual question answering
Researchers have developed ProMSA, a novel agent designed for knowledge-based visual question answering (KB-VQA). Unlike previous methods that use fixed retrieval pipelines, ProMSA adaptively selects between image searc…