PulseAugur
EN
LIVE 11:55:43
ENTITY multimodal large language model

multimodal large language model

PulseAugur coverage of multimodal large language model — every cluster mentioning multimodal large language model across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
37
120 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
37
119 over 90d
TIER MIX · 90D
TOPICS
RELATIONSHIPS
SENTIMENT · 30D

18 day(s) with sentiment data

RECENT · PAGE 1/6 · 120 TOTAL
  1. TOOL · CL_196045 ·

    VDC-Agent framework enables autonomous self-evolving video captioning

    Researchers have developed VDC-Agent, a novel framework that enables a single multimodal large language model to autonomously generate and refine video detailed captions. This self-evolving system overcomes the reliance…

  2. TOOL · CL_193650 ·

    New benchmark and framework for assessing visual spatial aesthetics in AI-generated images

    Researchers have introduced SA-BENCH, a new benchmark designed to evaluate the visual spatial aesthetics of interior scenes, a domain previously underserved by existing Image Quality Assessment (IQA) methods. The benchm…

  3. TOOL · CL_193266 ·

    MetaSpace framework tests spatial cognition in embodied AI agents

    Researchers have introduced MetaSpace, a novel framework for evaluating the spatial cognition of embodied agents. This system applies metamorphic testing principles, commonly used in software engineering, to automatical…

  4. TOOL · CL_187524 ·

    SEED system offers explainable detection for AI-generated text forgeries

    Researchers have developed SEED, a system designed to detect and explain AI-generated text forgeries. This system, which ranked third in the GenText-Forensics Challenge at ACM MM 2026, utilizes a Vision Transformer (ViT…

  5. TOOL · CL_187349 ·

    New text steganography uses dynamic codebook and multimodal LLM

    Researchers have developed a novel black-box text steganography method that enhances security and practicality by employing a dynamic codebook and a multimodal large language model. This approach addresses limitations o…

  6. TOOL · CL_187293 ·

    UniVVT framework uses multimodal LLM for end-to-end video virtual try-on

    Researchers have introduced UniVVT, a novel end-to-end framework for high-fidelity video virtual try-on. Unlike previous methods that rely on separate modules for human parsing, pose estimation, and garment warping, Uni…

  7. TOOL · CL_185383 ·

    New framework enhances spatial reasoning in multimodal LLMs without retraining

    Researchers have developed a new training-free framework designed to improve spatial reasoning in multimodal large language models (MLLMs). This framework, called Trace, Verify, and Correct, constructs a Spatial Evidenc…

  8. TOOL · CL_185336 ·

    New framework uses MLLMs for physically plausible video object insertion

    Researchers have developed Place-it-R1, a new framework designed to improve video object insertion by incorporating environment-aware reasoning from multimodal large language models (MLLMs). This approach ensures that i…

  9. TOOL · CL_185280 ·

    New MCU method improves continual unlearning in multimodal LLMs

    Researchers have developed a new method called Merging for Continual Unlearning (MCU) to address the challenges of removing specific information from multimodal large language models (MLLMs) without degrading their over…

  10. RESEARCH · CL_187161 ·

    PaDoc parser enables parallel document analysis, boosting speed and accuracy

    Researchers have developed PaDoc, a novel layout-grounded parser designed to improve the efficiency of document parsing. Unlike traditional end-to-end parsers that serialize content sequentially, PaDoc treats the docume…

  11. TOOL · CL_181045 ·

    New MIEScore model and MIE-Bench dataset advance multi-source image editing evaluation

    Researchers have introduced MIEScore, a new evaluation model designed to assess multi-source image editing (MIE) capabilities, which are crucial for advanced image manipulation tasks. Existing benchmarks often fall shor…

  12. TOOL · CL_180971 ·

    CoT-Edit framework enhances instruction-based video editing

    Researchers have introduced CoT-Edit, a novel framework for instruction-based video editing that addresses challenges in complex scenes. The system utilizes a Chain-of-Thought (CoT) enhanced multimodal large language mo…

  13. TOOL · CL_180561 ·

    New benchmark CultureVidBench assesses cultural understanding in text-to-video models

    Researchers have introduced CultureVidBench, a new benchmark designed to evaluate the cultural understanding capabilities of text-to-video generation models. This benchmark includes 1,000 prompts spanning 12 countries a…

  14. TOOL · CL_180553 ·

    Slot2Text introduces efficient object-centric visual tokenization for surgical MLLMs

    Researchers have developed Slot2Text, a novel approach for multimodal large language models (MLLMs) in surgical settings. This method replaces the typical dense visual tokens with a more efficient set of "slot latents" …

  15. RESEARCH · CL_183391 ·

    New STAMPlus model resolves segmentation trilemma for MLLMs

    Researchers have introduced STAMPlus, a novel approach to multimodal large language model (MLLM)-based segmentation that addresses the performance, dialogue ability, and inference speed trilemma. STAMPlus builds upon th…

  16. TOOL · CL_174282 ·

    New FAME benchmark standardizes evaluation for few-shot medical image segmentation

    Researchers have introduced FAME, a new benchmark designed to evaluate few-shot medical image segmentation (FS-MIS) methods. FAME standardizes evaluation across diverse approaches, including specialist models, SAM-based…

  17. RESEARCH · CL_174286 ·

    New SPaTS framework enhances MLLM scene text spotting capabilities

    Researchers have developed a new framework called Single-Patch Text Spotting (SPaTS) for Multimodal Large Language Models (MLLMs) to improve scene text spotting. This approach uses a single anchor visual token per text …

  18. TOOL · CL_169831 ·

    New MEDit-Bench dataset evaluates message-driven video editing

    Researchers have introduced MEDit-Bench, a new dataset designed to evaluate message-driven narrative video editing. This benchmark addresses the limitations of existing video summarization tasks by considering how diffe…

  19. RESEARCH · CL_173716 ·

    New OVEarth-Bench benchmark evaluates open-vocabulary Earth observation models

    A new benchmark called OVEarth-Bench has been introduced to evaluate open-vocabulary Earth observation capabilities. This benchmark addresses limitations in existing evaluations by expanding category breadth and query d…

  20. TOOL · CL_167801 ·

    New ConFusion framework enables fine-grained control over image fusion

    Researchers have developed ConFusion, a new framework for controllable infrared and visible image fusion. This method addresses limitations in existing approaches by learning a continuous fusion space, allowing for fine…