Large Multimodal Models
PulseAugur coverage of Large Multimodal Models — every cluster mentioning Large Multimodal Models across labs, papers, and developer communities, ranked by signal.
4 day(s) with sentiment data
-
New framework enhances AI reasoning for dense sports video analysis
Researchers have developed SportsGrounder, a new framework designed to improve the reasoning capabilities of Large Multimodal Models (LMMs) when analyzing dense sports videos. The framework addresses the challenge of di…
-
New ReGraph framework enables LMMs to generate structured recipe graphs from food images
Researchers have introduced ReGraph, a novel dataset and framework designed to enable Large Multimodal Models (LMMs) to generate structured recipe graphs from food images. This approach aims to explicitly represent ingr…
-
Open-source multimodal search agent POINTS-Seeker tackles context limits
Researchers have introduced POINTS-Seeker, an open-source multimodal search agent designed to overcome limitations in current Large Multimodal Models (LMMs). The system addresses the challenges of cultivating search age…
-
New framework uses LMMs to correct visual species recognition errors
A new research paper proposes a framework called Post-hoc Correction (POC) to improve visual species recognition (VSR) accuracy. The study found that while Large Multimodal Models (LMMs) underperform expert few-shot lea…
-
New VVM-Tuning framework enhances LMMs for unseen visual modalities
Researchers have developed a new training framework called VVM-Tuning to enhance the generalization capabilities of Large Multimodal Models (LMMs) across various visual modalities. This method synthesizes diverse visual…
-
New HART technique enables LMMs to reason with high-resolution images without annotations
Researchers have developed a new technique called HART (High-resolution Annotation-free Reasoning Technique) to improve how Large Multimodal Models (LMMs) handle high-resolution images. Current LMMs struggle with the la…
-
CarbonCLIP uses LMMs to improve satellite-based carbon emission prediction
Researchers have developed CarbonCLIP, a novel framework designed to enhance the accuracy of carbon emission predictions from satellite imagery. This approach integrates street-view semantics and temporal context, bridg…
-
CAIRN model advances multi-room 3D scene understanding
Researchers have introduced CAIRN, a novel topology-aware Large Multimodal Model designed for understanding complex multi-room 3D scenes. Unlike previous models that are limited to single rooms, CAIRN explicitly reasons…
-
New methods enhance multimodal industrial anomaly detection · 2 sources tracked
Researchers have developed two distinct methods for improving multimodal industrial anomaly detection. The first, Tuned Reverse Distillation (TRD), utilizes a multi-branch design and crossmodal tuners to enhance the lea…
-
New benchmarks and datasets advance deepfake detection for audio, image, and video
Researchers have introduced several new datasets and benchmarks aimed at improving the detection of deepfakes across various media. Echoes focuses on music deepfakes, emphasizing semantic alignment and provider diversit…
-
New Regularizer Enhances Taxonomic Knowledge in Large Multimodal Models
Researchers have developed a new method called Hierarchical Representation Regularization ($HiR^2$) to improve the taxonomic knowledge of large multimodal models (LMMs). Current LMMs often lack understanding of semantic…
-
New framework SAYRE synthesizes data to boost multimodal KIE models
Researchers have developed SAYRE, a novel framework for synthesizing training data to improve Key Information Extraction (KIE) capabilities in Large Multimodal Models (LMMs). This scene-aware synthesis approach generate…
-
SenseNova-Vision unifies computer vision tasks as multimodal generation · 6 sources tracked
Researchers have developed SenseNova-Vision, a unified multimodal model that treats all computer vision tasks as generation problems. This approach uses natural language instructions and visual prompts to specify tasks,…
-
New vLLM pipeline unifies audio generation and understanding
Researchers have developed a novel inference pipeline utilizing vLLM to unify audio understanding and generation tasks. This system addresses the challenges of high-throughput multimodal generation, particularly for spe…
-
New AI frameworks tackle long-form video understanding with advanced memory and reasoning
Researchers are developing advanced frameworks to improve how AI models understand and reason about long-form videos. Homer, for instance, uses a hierarchical memory system that organizes information by temporal and cau…
-
New HarmVideoBench evaluates LLMs on nuanced harmful video understanding · 2 sources tracked
Researchers have introduced HarmVideoBench, a new benchmark designed to evaluate the harmful video understanding capabilities of large vision-language models (LVLMs). Existing benchmarks often oversimplify harmful conte…
-
New LMM 'PreciseDoc' Enhances Document Element Grounding Accuracy
Researchers have developed PreciseDoc, a new Large Multimodal Model (LMM) designed to improve the accuracy of grounding specific elements within documents. Existing models struggle with precise localization in text-heav…
-
PRISMR framework enhances LMMs for multimodal listwise ranking
Researchers have developed PRISMR, a new framework designed to improve the performance of Large Multimodal Models (LMMs) in listwise ranking tasks, particularly in long-context scenarios. PRISMR addresses a failure mode…
-
New method uses location attention and LMMs for worldwide image geo-localization
Researchers have developed TransGeoCLIP, a new framework for worldwide image geo-localization that uses a location attention mechanism and large multimodal models. This method aims to improve accuracy by distinguishing …
-
Researchers isolate visual relation vectors in LMMs
Researchers have identified specific attention heads within Large Multimodal Models (LMMs) that are crucial for processing visual relations. By extracting and manipulating these "function vectors," they can improve the …