Multimodal Large Language Models and Tunings: Vision, Language, Sensors, Audio, and Beyond
PulseAugur coverage of Multimodal Large Language Models and Tunings: Vision, Language, Sensors, Audio, and Beyond — every cluster mentioning Multimodal Large Language Models and Tunings: Vision, Language, Sensors, Audio, and Beyond across labs, papers, and developer communities, ranked by signal.
- instance of ScienceCast 90%
- instance of CatalyzeX 90%
- instance of Gotit.pub 90%
- instance of CORE Recommender 90%
- instance of Qwen2.5-VL-7B 90%
- instance of Qwen2.5-VL 90%
- developed Multimodal Continual Instruction Tuning 90%
- instance of Qwen3-VL-8B-Instruct 90%
- instance of IU-Xray 90%
- instance of alphaXiv 70%
- used by Gotit.pub 70%
- used by ScienceCast 70%
13 day(s) with sentiment data
-
New survey maps Efficient Multimodal Learning landscape
A new survey paper systematically categorizes the field of Efficient Multimodal Learning (EML), addressing computational and memory bottlenecks in multimodal models. It proposes a model-to-system taxonomy, analyzing ove…
-
New method verifies object claims in multimodal LLMs
Researchers have developed a new training-free method called Semantic-Spatial Agreement Verification (SSAV) to address object hallucination in multimodal large language models. This technique verifies object claims by a…
-
MedSAM-3 enhances medical image segmentation with text prompts and LLM agents
Researchers have introduced MedSAM-3, a new model designed for medical image segmentation that leverages text prompts for precise targeting of anatomical structures. By fine-tuning the Segment Anything Model (SAM) archi…
-
New frameworks advance AI video understanding with structured reasoning and memory
Two new research papers introduce novel frameworks for enhancing video understanding and reasoning capabilities. The first, EventGraph and EventField, utilizes structured temporal representations to achieve high accurac…
-
New benchmarks assess MLLMs' geometric reasoning and visual perception
Researchers have developed new benchmarks to evaluate the geometric reasoning capabilities of Multimodal Large Language Models (MLLMs). The CapGeo-Bench, proposed in one study, uses high-quality figure-caption pairs and…
-
EMMI system enables efficient multimodal LLM inference on edge devices
Researchers have developed EMMI (Edge Multi-Modal Intelligence), a novel approach to make multimodal large language models (MLLMs) more efficient for edge devices. EMMI compresses multimodal representations at the edge …
-
New framework aligns multimodal LLMs with reasoning paths beyond imitation
Researchers have developed a new framework for multimodal in-context learning (ICL) that aims to improve how large language models (LLMs) align their responses with the reasoning process required by complex multimodal i…
-
VideoTIR method uses RL to improve long video understanding in LLMs
Researchers have developed VideoTIR, a novel method for improving the understanding of long videos by multimodal large language models (MLLMs). VideoTIR utilizes reinforcement learning to guide MLLMs in effectively usin…
-
New research tackles LLM jailbreaks with novel evaluation and defense methods · 4 sources tracked
Recent research papers explore novel methods for evaluating and defending against jailbreaking attempts on large language models (LLMs). One study systematically compares six automated jailbreak evaluators, finding that…
-
New Video Relational Algebra Optimizes Multimodal LLM Queries
Researchers have developed Concord, a system designed to optimize semantic video queries by introducing a Video Relational Algebra (VRA). This new algebra allows for operations on videos, transcripts, and object tracks,…
-
New benchmark and distillation methods advance on-device fire detection AI
Researchers are developing methods to compress large vision-language models (VLMs) for on-device deployment in safety-critical applications like fire detection. One approach involves a teacher-student knowledge distilla…
-
New MCPO method compresses multimodal LLM reasoning chains
Researchers have developed Modality-Contrastive Preference Optimization (MCPO), a novel method to compress lengthy reasoning chains in multimodal large language models. This technique addresses the computational costs a…
-
New VDiff-Bench benchmark reveals MLLMs struggle with subtle image differences
A new benchmark called VDiff-Bench has been introduced to evaluate the capabilities of multimodal large language models (MLLMs) in identifying subtle differences between images. The benchmark reveals significant weaknes…
-
New AI Model NeoRed Enhances Neonatal Respiratory Disease Diagnosis
Researchers have developed NeoRed, a novel multimodal large language model specifically designed for diagnosing neonatal respiratory diseases. This model addresses limitations in existing systems, such as the domain gap…
-
MARS framework improves text-video retrieval using multimodal LLMs
Researchers have developed MARS, a novel framework designed to enhance text-video retrieval by leveraging multiple layers and adaptive representation slots from multimodal large language models. Unlike existing methods …
-
New DocHop benchmark challenges MLLMs in multi-hop document reasoning
Researchers have introduced DocHop, a new benchmark designed to evaluate the multi-hop reasoning capabilities of Multimodal Large Language Models (MLLMs) when dealing with information-dense documents. Unlike existing be…
-
New benchmark CoMMET evaluates multimodal LLMs' Theory of Mind
Researchers have introduced CoMMET, a new benchmark designed to evaluate the Theory of Mind (ToM) capabilities of multimodal large language models (MLLMs). This benchmark is inspired by the psychology-based Theory of Mi…
-
New ClearText-Video dataset probes MLLM text-reading in low-quality videos
Researchers have introduced ClearText-Video (CTVid), a new large-scale dataset designed to evaluate how multimodal large language models (MLLMs) handle text-centric video understanding under varying quality conditions. …
-
Instruction Distillation enhances MLLM visual learning efficiency
Researchers have developed a new method called Instruction Distillation to improve the efficiency of visual in-context learning (ICL) in multimodal large language models (MLLMs). This offline procedure generates specifi…
-
New HalluPrism tool diagnoses MLLM failures with improved accuracy
Researchers have developed HalluPrism, a new diagnostic tool designed to better understand the failure modes of Multimodal Large Language Models (MLLMs). This method involves re-running model answers after introducing v…