PulseAugur
EN
LIVE 15:56:27
ENTITY Large Multimodal Models

Large Multimodal Models

PulseAugur coverage of Large Multimodal Models — every cluster mentioning Large Multimodal Models across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
44
44 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
43
43 over 90d
TIER MIX · 90D
TOPICS
SENTIMENT · 30D

3 day(s) with sentiment data

RECENT · PAGE 1/3 · 48 TOTAL
  1. TOOL · CL_275483 ·

    New benchmark UnifiedAttack targets LMM safety in harmful image-text generation

    Researchers have developed UnifiedAttack, a new benchmark to evaluate the safety of Large Multimodal Models (LMMs) in generating harmful content by coordinating text and image modalities. This approach aims to identify …

  2. RESEARCH · CL_256939 ·

    RegRet framework enhances region-level retrieval in Large Multimodal Models

    Researchers have introduced RegRet, a new framework designed to improve region-level retrieval in Large Multimodal Models (LMMs). This approach focuses on enhancing the capture of detailed regional features within image…

  3. RESEARCH · CL_245276 ·

    New LogiScope-VQA benchmark reveals LMMs lag human performance in industrial hazard identification

    A new benchmark dataset called LogiScope-VQA has been developed to evaluate the capabilities of large multimodal models (LMMs) in identifying logistics hazards within industrial settings. The dataset, comprising images,…

  4. TOOL · CL_229131 ·

    New benchmark tests AI's social reasoning in counterfactual videos

    Researchers have introduced SocialReasonBench, a new video question-answering benchmark designed to evaluate the social reasoning capabilities of Large Multimodal Models (LMMs). This benchmark utilizes counterfactual na…

  5. TOOL · CL_219046 ·

    New ImgCoder framework generates scientifically accurate images for AI reasoning

    A new research paper introduces ImgCoder, a framework designed to generate scientifically accurate images, addressing the limitations of current text-to-image models that often produce visually plausible but logically i…

  6. TOOL · CL_218266 ·

    New MedReaMM Benchmark Reveals LMMs Struggle with Clinical Diagnosis

    Researchers have introduced MedReaMM, a new benchmark designed to evaluate the diagnostic synthesis capabilities of Large Multimodal Models (LMMs) in clinical settings. Unlike previous benchmarks that focused on isolate…

  7. TOOL · CL_212193 ·

    ArmorOCR framework enhances adversarial OCR perception with new AdvSpot benchmark

    Researchers have introduced ArmorOCR, a novel two-stage training framework designed to enhance the robustness of optical character recognition (OCR) against adversarial attacks. This framework addresses the limitations …

  8. TOOL · CL_210614 ·

    New RL framework advances 3D point cloud quality assessment

    Researchers have introduced PCQA-R1, a novel reinforcement learning framework designed for no-reference 3D point cloud quality assessment. This system utilizes a chain-of-thought dataset and a Gaussian proximity reward …

  9. TOOL · CL_206164 ·

    New Latent-OPD method enhances LMMs for frame-efficient video reasoning

    Researchers have introduced Latent-OPD, a novel method for improving the efficiency of Large Multimodal Models (LMMs) in video reasoning. This technique enhances On-Policy Distillation (OPD) by incorporating trajectory-…

  10. TOOL · CL_193988 ·

    New framework enhances AI reasoning for dense sports video analysis

    Researchers have developed SportsGrounder, a new framework designed to improve the reasoning capabilities of Large Multimodal Models (LMMs) when analyzing dense sports videos. The framework addresses the challenge of di…

  11. TOOL · CL_191123 ·

    New ReGraph framework enables LMMs to generate structured recipe graphs from food images

    Researchers have introduced ReGraph, a novel dataset and framework designed to enable Large Multimodal Models (LMMs) to generate structured recipe graphs from food images. This approach aims to explicitly represent ingr…

  12. TOOL · CL_172069 ·

    Open-source multimodal search agent POINTS-Seeker tackles context limits

    Researchers have introduced POINTS-Seeker, an open-source multimodal search agent designed to overcome limitations in current Large Multimodal Models (LMMs). The system addresses the challenges of cultivating search age…

  13. TOOL · CL_143835 ·

    New framework uses LMMs to correct visual species recognition errors

    A new research paper proposes a framework called Post-hoc Correction (POC) to improve visual species recognition (VSR) accuracy. The study found that while Large Multimodal Models (LMMs) underperform expert few-shot lea…

  14. TOOL · CL_141746 ·

    New VVM-Tuning framework enhances LMMs for unseen visual modalities

    Researchers have developed a new training framework called VVM-Tuning to enhance the generalization capabilities of Large Multimodal Models (LMMs) across various visual modalities. This method synthesizes diverse visual…

  15. TOOL · CL_133657 ·

    New HART technique enables LMMs to reason with high-resolution images without annotations

    Researchers have developed a new technique called HART (High-resolution Annotation-free Reasoning Technique) to improve how Large Multimodal Models (LMMs) handle high-resolution images. Current LMMs struggle with the la…

  16. RESEARCH · CL_133135 ·

    CarbonCLIP uses LMMs to improve satellite-based carbon emission prediction

    Researchers have developed CarbonCLIP, a novel framework designed to enhance the accuracy of carbon emission predictions from satellite imagery. This approach integrates street-view semantics and temporal context, bridg…

  17. RESEARCH · CL_131399 ·

    CAIRN model advances multi-room 3D scene understanding

    Researchers have introduced CAIRN, a novel topology-aware Large Multimodal Model designed for understanding complex multi-room 3D scenes. Unlike previous models that are limited to single rooms, CAIRN explicitly reasons…

  18. RESEARCH · CL_129436 ·

    New methods enhance multimodal industrial anomaly detection · 2 sources tracked

    Researchers have developed two distinct methods for improving multimodal industrial anomaly detection. The first, Tuned Reverse Distillation (TRD), utilizes a multi-branch design and crossmodal tuners to enhance the lea…

  19. RESEARCH · CL_129070 ·

    New benchmarks and datasets advance deepfake detection for audio, image, and video

    Researchers have introduced several new datasets and benchmarks aimed at improving the detection of deepfakes across various media. Echoes focuses on music deepfakes, emphasizing semantic alignment and provider diversit…

  20. TOOL · CL_128765 ·

    New Regularizer Enhances Taxonomic Knowledge in Large Multimodal Models

    Researchers have developed a new method called Hierarchical Representation Regularization ($HiR^2$) to improve the taxonomic knowledge of large multimodal models (LMMs). Current LMMs often lack understanding of semantic…