multimodal large language model
PulseAugur coverage of multimodal large language model — every cluster mentioning multimodal large language model across labs, papers, and developer communities, ranked by signal.
- instance of CatalyzeX 90%
- instance of Multimodal Large Language Models and Tunings: Vision, Language, Sensors, Audio, and Beyond 90%
- used by ScienceCast 70%
- used by alphaXiv 70%
- used by CatalyzeX 70%
- instance of DagsHub 70%
- uses text-to-video generation 70%
- used by DagsHub 60%
- used by Gotit.pub 60%
- developed text-to-video generation 50%
- affiliated with text-to-video generation 50%
10 day(s) with sentiment data
-
New MIRAGE study reveals flaws in multimodal agent evidence use
A new study titled MIRAGE introduces a controlled evaluation framework for multimodal large language model (MLLM) agents, focusing on their ability to retrieve and utilize historical evidence across conversations. The r…
-
New RL method boosts MLLM visual perception with fewer tokens
Researchers have developed Vision-RL2, a novel reinforcement learning approach to enhance fine-grained visual perception in multimodal large language models (MLLMs). This method optimizes a region proposal network by tr…
-
New neuro-symbolic system MedVA improves medical volume visualization
Researchers have developed MedVA, a novel neuro-symbolic agentic system designed to enhance medical volume visualization. This system addresses limitations in current workflows by integrating three specialized agents: o…
-
DiVA simulator enables interactive digital life simulation via video models
Researchers have introduced DiVA, a novel digital life simulator designed for long-term, open-ended interactive experiences. DiVA utilizes a Multimodal Large Language Model (MLLM) as a router, coupled with a specialized…
-
New methods enhance MLLM understanding of long videos · 2 sources tracked
Two new research papers address the challenge of enabling multimodal large language models (MLLMs) to understand long videos, which is currently limited by token and computational budgets. The first paper, "Evaluation o…
-
New methods tackle test-time adaptation challenges in AI models · 2 sources tracked
Researchers have developed new methods for test-time adaptation in machine learning models. The first approach, MASA, uses a multimodal large language model to anchor semantic descriptions, helping to break a self-refer…
-
New framework STAG explains audio MLLM reasoning
Researchers have developed STAG, a novel post-hoc framework designed to explain the reasoning behind audio-based multimodal large language models (MLLMs). This system provides token-level spectro-temporal grounding, ide…
-
New framework enables efficient long-video understanding on edge devices
Researchers have developed a new framework called Caption-once, Frames-on-Demand (CFD) designed for efficient long-video understanding on edge devices. This system utilizes a dual-track narrative index, combining an eve…
-
Study finds flaws in MC-VQA benchmarks for LLM evaluation
A new study published on arXiv reveals significant flaws in multiple-choice Visual Question Answering (MC-VQA) benchmarks, which are commonly used to evaluate Multimodal Large Language Models (MLLMs). Researchers found …
-
New AI agent accurately converts blueprints to simulation models
Researchers have developed BlueprintAgent (BPA), a novel multimodal agent designed to accurately convert scanned structural blueprints into simulation-ready models. Unlike previous methods that relied on direct promptin…
-
New AI agent StrixAE enhances audio with multimodal LLM
Researchers have developed StrixAE, an intelligent agent designed for audio enhancement in complex real-world scenarios. StrixAE utilizes a multimodal large language model (MLLM) to manage various audio enhancement and …
-
AI image forensics research advances detection and localization techniques
Two research papers submitted to arXiv propose advanced methods for detecting and localizing manipulated images, particularly those generated by AI. The first paper introduces an evidence-guided system for the GenText-F…
-
New AI Agent Enhances Image Focus Using MLLM and Diffusion Retouching
Researchers have developed EyeControl, a new agent that uses a multimodal large language model (MLLM) and a diffusion-based retouching executor to enhance visual focus in images. This system allows users to guide attent…
-
New 'Distributed Implicit Harm' vulnerability found in MLLM video moderation
Researchers have identified a new safety vulnerability in multimodal large language models (MLLMs) used for video moderation, termed Distributed Implicit Harm (DIH). This occurs when seemingly harmless video components …
-
ShallowStream framework enhances MLLM streaming video understanding
Researchers have introduced ShallowStream, a new framework designed to improve the efficiency of streaming video understanding with multimodal large language models (MLLMs). By utilizing the shallow layers of an MLLM, S…
-
New TUE-Detector uses MLLMs to identify AI-generated videos
Researchers have developed TUE-Detector, a new framework designed to identify AI-generated videos by leveraging multimodal large language models (MLLMs). This system functions as a tool-using expert, trained to invoke s…
-
New methods enhance camouflaged object detection using multispectral and multimodal AI
Researchers have developed new methods for camouflaged object detection (COD) that go beyond traditional RGB imagery. One approach, MSFormer, utilizes multispectral images to capture richer spectral signatures, outperfo…
-
New benchmarks and frameworks advance multimodal AI agents
Researchers have introduced new frameworks and benchmarks to improve multimodal search agents. WeAgent-Harness and WeAgent-MMSearch aim to enable agents to natively interact with and cite images retrieved from the web, …
-
CAMIE framework enhances product ad retrieval with multimodal embeddings
Researchers have developed CAMIE, a novel framework for multimodal item embeddings designed to improve retrieval in dynamic product advertising systems. This framework leverages large language and multimodal models to r…
-
New research advances multi-object tracking with 3D geometry and LLM integration · 4 sources tracked
Researchers have developed new methods for multi-object tracking in videos, aiming to improve accuracy and efficiency. PLANET, a new end-to-end tracker, moves beyond image-plane limitations by incorporating 3D scene geo…