OK-VQA
PulseAugur coverage of OK-VQA — every cluster mentioning OK-VQA across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
New benchmark reveals faithfulness gaps in Vision-Language Models
Researchers have introduced EDCT-Bench, a new benchmark designed to identify faithfulness issues in Vision-Language Models (VLMs). This benchmark uses an intervention-based protocol called Explanation-Driven Counterfact…
-
New privacy defense prunes visual tokens for LLMs
Researchers have developed QPriv-VL, a novel framework designed to enhance privacy in Vision-Language Models (VLMs) used in sensitive applications like Federated Learning. This system intelligently prunes visual tokens …
-
GraphLoom framework improves multimodal RAG with knowledge graphs
Researchers have introduced GraphLoom, a novel framework designed to enhance multimodal retrieval-augmented generation (RAG) systems. This system constructs a multimodal knowledge graph from various data sources, includ…
-
New SKIP Architecture Slashes Multimodal QA Costs with Sparse Routing
Researchers have introduced SKIP, a novel architecture for knowledge-intensive multimodal question answering that significantly reduces computational costs. SKIP achieves this by routing computation along sparse pathway…
-
New framework enhances MLLM knowledge reasoning for visual question answering
Researchers have developed a new framework called Hindsight Distilled Reasoning (HinD) to improve the knowledge reasoning capabilities of multimodal large language models (MLLMs) in visual question answering tasks. The …
-
New MAD-RAG method tackles Attention Distraction in LVLMs
Researchers have identified a new failure mode in retrieval-augmented large vision-language models (LVLMs) called Attention Distraction (AD). This occurs when highly relevant retrieved text globally suppresses visual at…