V$^*Bench
PulseAugur coverage of V$^*Bench — every cluster mentioning V$^*Bench across labs, papers, and developer communities, ranked by signal.
3 day(s) with sentiment data
-
New GAM-Agent framework boosts visual reasoning in LLMs via game theory
Researchers have developed GAM-Agent, a novel framework that enhances visual reasoning in large language models by employing a game-theoretic approach. This system treats the reasoning process as a non-zero-sum game whe…
-
New benchmarks and methods advance multimodal reasoning in AI
Researchers are developing new methods for multimodal knowledge graph completion and reasoning, integrating vision-language models (VLMs) with graph structures. ViSR-KGC proposes a visual subgraph reasoning approach tha…
-
New 'Thinking-Once' method improves high-resolution VQA by routing existing evidence
Researchers have developed a new method called Thinking-Once for high-resolution visual question answering (HR-VQA). This technique focuses on efficiently routing evidence that is already present in intermediate layers …
-
New HART technique enables LMMs to reason with high-resolution images without annotations
Researchers have developed a new technique called HART (High-resolution Annotation-free Reasoning Technique) to improve how Large Multimodal Models (LMMs) handle high-resolution images. Current LMMs struggle with the la…
-
New methods accelerate agentic LLM inference with speculative execution · 2 sources tracked
Two research papers introduce novel methods to accelerate the inference speed of agentic large language models (LLMs) by employing speculative execution. The first paper, SPORK, utilizes a lightweight probe from the LLM…
-
ActiveScope framework enhances MLLM perception by correcting errors
Researchers have introduced ActiveScope, a novel training-free framework designed to improve the perception capabilities of Multimodal Large Language Models (MLLMs). This framework addresses limitations in high-resoluti…
-
HiDe framework boosts MLLM performance on high-res images
Researchers have developed a new training-free framework called HiDe to improve the performance of Multimodal Large Language Models (MLLMs) on high-resolution images. HiDe addresses background interference rather than o…
-
Perceptual Flow Network and VGR enhance visual reasoning in LLMs
Researchers have developed a Perceptual Flow Network (PFlowNet) to improve visual reasoning in Large-Vision Language Models (LVLMs). PFlowNet decouples perception from reasoning and uses variational reinforcement learni…
-
SIEVES method boosts multimodal LLM coverage on visual tasks with evidence scoring
Researchers have developed SIEVES, a novel method for improving the reliability of multimodal large language models (MLLMs) in out-of-distribution scenarios. SIEVES works by learning to estimate the quality of visual ev…