MLLMs
PulseAugur coverage of MLLMs — every cluster mentioning MLLMs across labs, papers, and developer communities, ranked by signal.
- instance of multimodal large language model 95%
- instance of Multimodal Large Language Models and Tunings: Vision, Language, Sensors, Audio, and Beyond 90%
- instance of DagsHub 90%
- instance of alphaXiv 90%
- instance of CatalyzeX 90%
- instance of Gotit.pub 90%
- instance of Multimodal LLMs 90%
- instance of Qwen2.5-VL 90%
- developed Multimodal Large Language Models and Tunings: Vision, Language, Sensors, Audio, and Beyond 70%
- developed Gotit.pub 70%
- developed by Gotit.pub 70%
- developed by alphaXiv 70%
- 2026-05-22 research_milestone A new pipeline was introduced to enhance MLLMs for safety-critical driving video analysis. source
- 2026-05-22 research_milestone Researchers reveal and propose a method to recover temporal grounding in multimodal large language models. source
- 2026-05-22 research_milestone A new benchmark and dataset were introduced to evaluate MLLMs' ability to reason about personality beyond superficial cues. source
- 2026-05-21 research_milestone A new method using MLLMs for detecting AI-generated Chinese poetry achieves state-of-the-art results. source
21 day(s) with sentiment data
-
New benchmark and model advance camera motion understanding in videos
Researchers have introduced CamChoreo, a new benchmark dataset designed for understanding complex camera motions in videos. This dataset features 4,229 real-world clips with detailed temporal annotations, where nearly h…
-
New framework evaluates MLLMs for thermal grounding in infrared images
A new evaluation framework for multimodal large language models (MLLMs) applied to infrared images has been developed, which goes beyond simple answer accuracy to assess the grounding of explanations in thermal evidence…
-
New framework enhances multimodal LLM reasoning for visual spatial intelligence
Researchers have introduced an "Advantage-Guided Gate" framework to improve the open-ended reasoning capabilities of multimodal large language models (MLLMs) in visual spatial intelligence tasks. This framework addresse…
-
New VCU-Bridge framework enhances MLLM visual reasoning hierarchy
Researchers have introduced VCU-Bridge, a new framework designed to improve how Multimodal Large Language Models (MLLMs) understand visual information. Unlike current models that often process details and high-level con…
-
M3Prune framework enhances multi-modal AI agent efficiency by pruning communication
A new framework called M$^3$Prune has been developed to improve the efficiency of multi-modal retrieval-augmented generation (mRAG) systems. This framework addresses the high token overhead and computational costs assoc…
-
New AI methods enhance GUI grounding with self-evolution and reflection · 4 sources tracked
Researchers are developing advanced methods for GUI visual grounding, enabling AI agents to better interact with graphical user interfaces. One approach, Test-Time Self-Evolving GUI Visual Grounding, uses a closed-loop …
-
New AI Frameworks Tackle Visual Token Pruning in Multimodal LLMs
Researchers are developing new methods to optimize multimodal large language models (MLLMs) by pruning visual tokens, which are computationally expensive. One approach, MAP, predicts the importance of visual tokens by l…
-
New CASA method boosts multimodal LLM safety alignment
Researchers have developed CASA (Classification Augmented with Safety Attention), a novel strategy to enhance the safety alignment of multimodal large-language models (MLLMs). CASA uses internal MLLM representations to …
-
New benchmark tests MLLMs for proactive safety, revealing significant gaps
Researchers have developed SPRINT, a new benchmark designed to evaluate the proactive risk inference capabilities of multimodal large language models (MLLMs). The benchmark utilizes 2,888 real-world sports videos, inclu…
-
New MM-ISTS framework uses multimodal LLMs for irregular time series forecasting
Researchers have developed MM-ISTS, a novel framework designed to improve the forecasting of irregularly sampled time series data. This approach integrates multimodal vision-text large language models (MLLMs) to capture…
-
New RL strategy trains MLLMs to refuse non-existent objects
Researchers have developed a new reinforcement learning strategy called Refusal-Calibrated Group Relative Policy Optimization (RC-GRPO) to improve the ability of Multimodal Large Language Models (MLLMs) to correctly ide…
-
New VTS method reveals MLLMs struggle with visual instructions
Researchers have introduced Visualized Task Semantics (VTS), a novel method to evaluate how well multimodal large language models (MLLMs) understand instructions presented visually within images. When questions were mov…
-
New AI models aim for deeper emotional understanding and reasoning · 5 sources tracked
Researchers are developing advanced multimodal AI models capable of understanding and reasoning about human emotions. Several new papers introduce frameworks and benchmarks for this purpose, focusing on integrating verb…
-
New C4 framework evaluates MLLM creativity using Chinese idioms · 2 sources tracked
Researchers have introduced C4, a new evaluation framework designed to assess the cross-concept creativity of Multimodal Large Language Models (MLLMs). This framework utilizes Chinese idioms (Chengyu) to test a model's …
-
New ChartAnno Benchmark Evaluates MLLMs on Chart Annotation Generation
A new benchmark called ChartAnno has been developed to evaluate how well multimodal large language models (MLLMs) can generate annotations for existing charts. The benchmark includes 1,200 real-world charts with associa…
-
New framework PixVL enhances pixel-level multimodal LLMs with self-supervision
Researchers have developed PixVL, a novel self-supervised post-training framework designed to enhance pixel-level multimodal large language models (MLLMs). This framework addresses the scarcity of labeled mask-text pair…
-
New MCR-GRPO framework enhances MLLMs for structured visual perception tasks
Researchers have introduced MCR-GRPO, a novel framework designed to improve the performance of Multimodal Large Language Models (MLLMs) on structured visual perception tasks. This new method addresses the granularity mi…
-
ET-Prune framework optimizes MLLM inference by dynamically pruning visual tokens
Researchers have developed ET-Prune, a novel framework designed to optimize the inference costs of multimodal large language models (MLLMs) by dynamically pruning visual tokens. Unlike fixed pruning ratios, ET-Prune ada…
-
New benchmarks and methods enhance MLLM chart reasoning capabilities
Researchers have developed new benchmarks and methods to improve the reasoning capabilities of multimodal large language models (MLLMs) when analyzing complex charts. The LongChart VQA benchmark introduces a dataset wit…
-
New benchmarks reveal MLLMs struggle with multi-step spatial reasoning
A new benchmark called LEGO-Puzzles has been developed to assess the multi-step spatial reasoning capabilities of Multimodal Large Language Models (MLLMs). Researchers found that even the most advanced MLLMs struggle wi…