multimodal large language model
PulseAugur coverage of multimodal large language model — every cluster mentioning multimodal large language model across labs, papers, and developer communities, ranked by signal.
18 day(s) with sentiment data
-
VDC-Agent framework enables autonomous self-evolving video captioning
Researchers have developed VDC-Agent, a novel framework that enables a single multimodal large language model to autonomously generate and refine video detailed captions. This self-evolving system overcomes the reliance…
-
New benchmark and framework for assessing visual spatial aesthetics in AI-generated images
Researchers have introduced SA-BENCH, a new benchmark designed to evaluate the visual spatial aesthetics of interior scenes, a domain previously underserved by existing Image Quality Assessment (IQA) methods. The benchm…
-
MetaSpace framework tests spatial cognition in embodied AI agents
Researchers have introduced MetaSpace, a novel framework for evaluating the spatial cognition of embodied agents. This system applies metamorphic testing principles, commonly used in software engineering, to automatical…
-
SEED system offers explainable detection for AI-generated text forgeries
Researchers have developed SEED, a system designed to detect and explain AI-generated text forgeries. This system, which ranked third in the GenText-Forensics Challenge at ACM MM 2026, utilizes a Vision Transformer (ViT…
-
New text steganography uses dynamic codebook and multimodal LLM
Researchers have developed a novel black-box text steganography method that enhances security and practicality by employing a dynamic codebook and a multimodal large language model. This approach addresses limitations o…
-
UniVVT framework uses multimodal LLM for end-to-end video virtual try-on
Researchers have introduced UniVVT, a novel end-to-end framework for high-fidelity video virtual try-on. Unlike previous methods that rely on separate modules for human parsing, pose estimation, and garment warping, Uni…
-
New framework enhances spatial reasoning in multimodal LLMs without retraining
Researchers have developed a new training-free framework designed to improve spatial reasoning in multimodal large language models (MLLMs). This framework, called Trace, Verify, and Correct, constructs a Spatial Evidenc…
-
New framework uses MLLMs for physically plausible video object insertion
Researchers have developed Place-it-R1, a new framework designed to improve video object insertion by incorporating environment-aware reasoning from multimodal large language models (MLLMs). This approach ensures that i…
-
New MCU method improves continual unlearning in multimodal LLMs
Researchers have developed a new method called Merging for Continual Unlearning (MCU) to address the challenges of removing specific information from multimodal large language models (MLLMs) without degrading their over…
-
PaDoc parser enables parallel document analysis, boosting speed and accuracy
Researchers have developed PaDoc, a novel layout-grounded parser designed to improve the efficiency of document parsing. Unlike traditional end-to-end parsers that serialize content sequentially, PaDoc treats the docume…
-
New MIEScore model and MIE-Bench dataset advance multi-source image editing evaluation
Researchers have introduced MIEScore, a new evaluation model designed to assess multi-source image editing (MIE) capabilities, which are crucial for advanced image manipulation tasks. Existing benchmarks often fall shor…
-
CoT-Edit framework enhances instruction-based video editing
Researchers have introduced CoT-Edit, a novel framework for instruction-based video editing that addresses challenges in complex scenes. The system utilizes a Chain-of-Thought (CoT) enhanced multimodal large language mo…
-
New benchmark CultureVidBench assesses cultural understanding in text-to-video models
Researchers have introduced CultureVidBench, a new benchmark designed to evaluate the cultural understanding capabilities of text-to-video generation models. This benchmark includes 1,000 prompts spanning 12 countries a…
-
Slot2Text introduces efficient object-centric visual tokenization for surgical MLLMs
Researchers have developed Slot2Text, a novel approach for multimodal large language models (MLLMs) in surgical settings. This method replaces the typical dense visual tokens with a more efficient set of "slot latents" …
-
New STAMPlus model resolves segmentation trilemma for MLLMs
Researchers have introduced STAMPlus, a novel approach to multimodal large language model (MLLM)-based segmentation that addresses the performance, dialogue ability, and inference speed trilemma. STAMPlus builds upon th…
-
New FAME benchmark standardizes evaluation for few-shot medical image segmentation
Researchers have introduced FAME, a new benchmark designed to evaluate few-shot medical image segmentation (FS-MIS) methods. FAME standardizes evaluation across diverse approaches, including specialist models, SAM-based…
-
New SPaTS framework enhances MLLM scene text spotting capabilities
Researchers have developed a new framework called Single-Patch Text Spotting (SPaTS) for Multimodal Large Language Models (MLLMs) to improve scene text spotting. This approach uses a single anchor visual token per text …
-
New MEDit-Bench dataset evaluates message-driven video editing
Researchers have introduced MEDit-Bench, a new dataset designed to evaluate message-driven narrative video editing. This benchmark addresses the limitations of existing video summarization tasks by considering how diffe…
-
New OVEarth-Bench benchmark evaluates open-vocabulary Earth observation models
A new benchmark called OVEarth-Bench has been introduced to evaluate open-vocabulary Earth observation capabilities. This benchmark addresses limitations in existing evaluations by expanding category breadth and query d…
-
New ConFusion framework enables fine-grained control over image fusion
Researchers have developed ConFusion, a new framework for controllable infrared and visible image fusion. This method addresses limitations in existing approaches by learning a continuous fusion space, allowing for fine…