Large Multimodal Models
PulseAugur coverage of Large Multimodal Models — every cluster mentioning Large Multimodal Models across labs, papers, and developer communities, ranked by signal.
3 day(s) with sentiment data
-
New benchmark UnifiedAttack targets LMM safety in harmful image-text generation
Researchers have developed UnifiedAttack, a new benchmark to evaluate the safety of Large Multimodal Models (LMMs) in generating harmful content by coordinating text and image modalities. This approach aims to identify …
-
RegRet framework enhances region-level retrieval in Large Multimodal Models
Researchers have introduced RegRet, a new framework designed to improve region-level retrieval in Large Multimodal Models (LMMs). This approach focuses on enhancing the capture of detailed regional features within image…
-
New LogiScope-VQA benchmark reveals LMMs lag human performance in industrial hazard identification
A new benchmark dataset called LogiScope-VQA has been developed to evaluate the capabilities of large multimodal models (LMMs) in identifying logistics hazards within industrial settings. The dataset, comprising images,…
-
New benchmark tests AI's social reasoning in counterfactual videos
Researchers have introduced SocialReasonBench, a new video question-answering benchmark designed to evaluate the social reasoning capabilities of Large Multimodal Models (LMMs). This benchmark utilizes counterfactual na…
-
New ImgCoder framework generates scientifically accurate images for AI reasoning
A new research paper introduces ImgCoder, a framework designed to generate scientifically accurate images, addressing the limitations of current text-to-image models that often produce visually plausible but logically i…
-
New MedReaMM Benchmark Reveals LMMs Struggle with Clinical Diagnosis
Researchers have introduced MedReaMM, a new benchmark designed to evaluate the diagnostic synthesis capabilities of Large Multimodal Models (LMMs) in clinical settings. Unlike previous benchmarks that focused on isolate…
-
ArmorOCR framework enhances adversarial OCR perception with new AdvSpot benchmark
Researchers have introduced ArmorOCR, a novel two-stage training framework designed to enhance the robustness of optical character recognition (OCR) against adversarial attacks. This framework addresses the limitations …
-
New RL framework advances 3D point cloud quality assessment
Researchers have introduced PCQA-R1, a novel reinforcement learning framework designed for no-reference 3D point cloud quality assessment. This system utilizes a chain-of-thought dataset and a Gaussian proximity reward …
-
New Latent-OPD method enhances LMMs for frame-efficient video reasoning
Researchers have introduced Latent-OPD, a novel method for improving the efficiency of Large Multimodal Models (LMMs) in video reasoning. This technique enhances On-Policy Distillation (OPD) by incorporating trajectory-…
-
New framework enhances AI reasoning for dense sports video analysis
Researchers have developed SportsGrounder, a new framework designed to improve the reasoning capabilities of Large Multimodal Models (LMMs) when analyzing dense sports videos. The framework addresses the challenge of di…
-
New ReGraph framework enables LMMs to generate structured recipe graphs from food images
Researchers have introduced ReGraph, a novel dataset and framework designed to enable Large Multimodal Models (LMMs) to generate structured recipe graphs from food images. This approach aims to explicitly represent ingr…
-
Open-source multimodal search agent POINTS-Seeker tackles context limits
Researchers have introduced POINTS-Seeker, an open-source multimodal search agent designed to overcome limitations in current Large Multimodal Models (LMMs). The system addresses the challenges of cultivating search age…
-
New framework uses LMMs to correct visual species recognition errors
A new research paper proposes a framework called Post-hoc Correction (POC) to improve visual species recognition (VSR) accuracy. The study found that while Large Multimodal Models (LMMs) underperform expert few-shot lea…
-
New VVM-Tuning framework enhances LMMs for unseen visual modalities
Researchers have developed a new training framework called VVM-Tuning to enhance the generalization capabilities of Large Multimodal Models (LMMs) across various visual modalities. This method synthesizes diverse visual…
-
New HART technique enables LMMs to reason with high-resolution images without annotations
Researchers have developed a new technique called HART (High-resolution Annotation-free Reasoning Technique) to improve how Large Multimodal Models (LMMs) handle high-resolution images. Current LMMs struggle with the la…
-
CarbonCLIP uses LMMs to improve satellite-based carbon emission prediction
Researchers have developed CarbonCLIP, a novel framework designed to enhance the accuracy of carbon emission predictions from satellite imagery. This approach integrates street-view semantics and temporal context, bridg…
-
CAIRN model advances multi-room 3D scene understanding
Researchers have introduced CAIRN, a novel topology-aware Large Multimodal Model designed for understanding complex multi-room 3D scenes. Unlike previous models that are limited to single rooms, CAIRN explicitly reasons…
-
New methods enhance multimodal industrial anomaly detection · 2 sources tracked
Researchers have developed two distinct methods for improving multimodal industrial anomaly detection. The first, Tuned Reverse Distillation (TRD), utilizes a multi-branch design and crossmodal tuners to enhance the lea…
-
New benchmarks and datasets advance deepfake detection for audio, image, and video
Researchers have introduced several new datasets and benchmarks aimed at improving the detection of deepfakes across various media. Echoes focuses on music deepfakes, emphasizing semantic alignment and provider diversit…
-
New Regularizer Enhances Taxonomic Knowledge in Large Multimodal Models
Researchers have developed a new method called Hierarchical Representation Regularization ($HiR^2$) to improve the taxonomic knowledge of large multimodal models (LMMs). Current LMMs often lack understanding of semantic…