MLLMs
PulseAugur coverage of MLLMs — every cluster mentioning MLLMs across labs, papers, and developer communities, ranked by signal.
- instance of Multimodal LLMs 95%
- instance of multimodal large language model 95%
- instance of alphaXiv 90%
- instance of Gotit.pub 90%
- instance of CatalyzeX 90%
- instance of Multimodal Large Language Models and Tunings: Vision, Language, Sensors, Audio, and Beyond 90%
- instance of DagsHub 90%
- instance of Qwen2.5-VL 90%
- developed Multimodal Large Language Models and Tunings: Vision, Language, Sensors, Audio, and Beyond 70%
- used by DagsHub 70%
- used by alphaXiv 70%
- used by Gotit.pub 70%
- 2026-05-22 research_milestone A new pipeline was introduced to enhance MLLMs for safety-critical driving video analysis. source
- 2026-05-22 research_milestone Researchers reveal and propose a method to recover temporal grounding in multimodal large language models. source
- 2026-05-22 research_milestone A new benchmark and dataset were introduced to evaluate MLLMs' ability to reason about personality beyond superficial cues. source
- 2026-05-21 research_milestone A new method using MLLMs for detecting AI-generated Chinese poetry achieves state-of-the-art results. source
11 day(s) with sentiment data
-
New methods enhance MLLM efficiency for long video analysis · 3 sources tracked
Researchers are developing new methods to improve the efficiency and accuracy of multimodal large language models (MLLMs) when processing long videos. VideoMM proposes an adaptive approach that separates semantic filter…
-
New adapter enables backward compatibility for multi-modal LLMs without model updates
Researchers have developed the Multi-modal Knowledge Preserving Adapter (MKP-Adapter), a novel approach for Multi-modal Large Language Models (MLLMs) that enables backward compatibility without updating the core model. …
-
New benchmark reveals MLLMs struggle with egocentric puzzle assistance
Researchers have developed PuzzleMate, a new framework and benchmark designed to evaluate the capabilities of Multimodal Large Language Models (MLLMs) in providing step-by-step guidance for complex physical tasks, using…
-
New CORAL framework enhances medical report generation with interpretable reasoning
Researchers have developed CORAL, a novel multimodal framework designed to improve the interpretability and accuracy of medical report generation from imaging data. This framework integrates spatial grounding and concep…
-
New methods accelerate Vision-Language Model inference by optimizing token processing · 2 sources tracked
Two new research papers propose methods to accelerate the inference of Vision-Language Models (VLMs) by reducing computational overhead. StackTok focuses on adaptive visual token selection, prioritizing query relevance …
-
New methods enhance MLLM understanding of long videos · 2 sources tracked
Two new research papers address the challenge of enabling multimodal large language models (MLLMs) to understand long videos, which is currently limited by token and computational budgets. The first paper, "Evaluation o…
-
New RL Framework Boosts MLLM Visual Math Reasoning
Researchers have introduced UniCAR-RL, a novel reinforcement learning framework designed to enhance the visual mathematical reasoning capabilities of Multimodal Large Language Models (MLLMs). This framework addresses th…
-
New benchmarks assess MLLMs' geometric reasoning and visual perception
Researchers have developed new benchmarks to evaluate the geometric reasoning capabilities of Multimodal Large Language Models (MLLMs). The CapGeo-Bench, proposed in one study, uses high-quality figure-caption pairs and…
-
New VIP-Router optimizes MLLM vision token pruning with adaptive strategy selection
Researchers have developed VIP-Router, a novel system designed to optimize the efficiency of multimodal large language models (MLLMs) by adaptively selecting the best vision token pruning strategy for each input. Unlike…
-
AI and LLMs integrated with TCAD for semiconductor design
A new paper explores the integration of AI, including machine learning and large language models (LLMs), with Technology Computer-Aided Design (TCAD) for semiconductor device design and defect discovery. The research de…
-
New HEAL method tackles MLLM hallucinations by calibrating information distribution
Researchers have developed a new method called HEAL to address hallucinations in Multimodal Large Language Models (MLLMs). HEAL identifies and mitigates hallucinations by analyzing information distribution within the mo…
-
New MV-STRIDE dataset boosts MLLM spatial reasoning capabilities
Researchers have introduced MV-STRIDE, a novel dataset designed to enhance the multi-view spatial reasoning capabilities of Multimodal Large Language Models (MLLMs). This dataset addresses a key limitation in current ML…
-
New CrossModalQA benchmark tests multimodal LLMs on complex reasoning
Researchers have introduced CrossModalQA, a new benchmark designed to evaluate multimodal large language models (MLLMs) in retrieval-augmented generation (RAG) tasks. This benchmark addresses limitations in existing sys…
-
New benchmark and distillation methods advance on-device fire detection AI
Researchers are developing methods to compress large vision-language models (VLMs) for on-device deployment in safety-critical applications like fire detection. One approach involves a teacher-student knowledge distilla…
-
New ReactHuman benchmark tests LLMs for robot reactive safety
Researchers have introduced ReactHuman, a new benchmark designed to test the reactive decision-making capabilities of multimodal large language models (MLLMs) in simulated humanoid robots. The benchmark focuses on immed…
-
New S^3-Bench framework evaluates MLLMs as scientific voice assistants
A new evaluation framework called S$^3$-Bench has been introduced to assess the capabilities of multimodal large language models (MLLMs) specifically as scientific voice assistants. This framework addresses the challeng…
-
ReactVAU framework enables real-time video anomaly understanding
Researchers have introduced ReactVAU, a novel framework designed for real-time video anomaly understanding in streaming environments. This system employs a dual-module approach, featuring a lightweight Fast Detection Mo…
-
New VDiff-Bench benchmark reveals MLLMs struggle with subtle image differences
A new benchmark called VDiff-Bench has been introduced to evaluate the capabilities of multimodal large language models (MLLMs) in identifying subtle differences between images. The benchmark reveals significant weaknes…
-
New benchmark InSituMeasure reveals MLLMs struggle with industrial measurement tasks
Researchers have introduced InSituMeasure, a new benchmark designed to evaluate the situated measurement grounding capabilities of Multimodal Large Language Models (MLLMs) in industrial settings. The benchmark includes …
-
New DocHop benchmark challenges MLLMs in multi-hop document reasoning
Researchers have introduced DocHop, a new benchmark designed to evaluate the multi-hop reasoning capabilities of Multimodal Large Language Models (MLLMs) when dealing with information-dense documents. Unlike existing be…