PulseAugur
EN
LIVE 22:04:19
ENTITY LVLMs

LVLMs

PulseAugur coverage of LVLMs — every cluster mentioning LVLMs across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
7
30 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
7
30 over 90d
TIER MIX · 90D
TOPICS
RELATIONSHIPS
SENTIMENT · 30D

4 day(s) with sentiment data

RECENT · PAGE 1/3 · 47 TOTAL
  1. TOOL · CL_254505 ·

    New TestHallVQA benchmark probes LVLMs' reasoning with redundant context

    Researchers have introduced TestHallVQA, a new benchmark designed to evaluate Large Vision-Language Models (LVLMs) on their ability to perform document-level reasoning, particularly in the presence of redundant or irrel…

  2. TOOL · CL_239296 ·

    New PetQA benchmark evaluates AI veterinary knowledge

    Researchers have developed PetQA, a new benchmark designed to evaluate the veterinary knowledge and clinical reasoning capabilities of large language models (LLMs) and large vision-language models (LVLMs). The benchmark…

  3. RESEARCH · CL_245016 ·

    New synthetic dataset enhances LVLM spatial reasoning capabilities

    Researchers have introduced SpatialBlock, a synthetic dataset designed to improve the spatial intelligence of Large Vision-Language Models (LVLMs). The dataset comprises 15,000 block-stacking problems that cover 3D-to-2…

  4. TOOL · CL_228882 ·

    New Cen-Prune method enhances LVLM efficiency by optimizing visual token pruning

    Researchers have developed a new method called Cen-Prune to improve the efficiency of Large Vision-Language Models (LVLMs) by optimizing how visual tokens are pruned. Standard diversity-based pruning relies on cosine si…

  5. TOOL · CL_227064 ·

    New dataset Synth-JDoc boosts LVLM Japanese OCR performance

    Researchers have developed Synth-JDoc, a novel dataset designed to improve the optical character recognition (OCR) capabilities of Large Vision Language Models (LVLMs), particularly for Japanese text. This dataset synth…

  6. TOOL · CL_218893 ·

    New benchmark reveals LVLMs struggle with interactive visual grounding

    A new research paper introduces a framework for evaluating interactive visual grounding in large vision-language models (LVLMs). The study highlights that current LVLMs struggle with tasks requiring dialogue to refine o…

  7. TOOL · CL_218209 ·

    New benchmark and training method boost multilingual LVLM performance

    Researchers have introduced PM4Bench, a new benchmark designed to evaluate the multilingual capabilities of Large Vision-Language Models (LVLMs). This benchmark utilizes a parallel corpus across 10 languages and incorpo…

  8. RESEARCH · CL_218205 ·

    New research tackles AI hallucinations in video and language models

    Researchers are developing new methods to combat hallucinations in AI models, particularly in video-language and large language models. One approach, CounterVid, uses counterfactual video generation to create synthetic …

  9. RESEARCH · CL_212014 ·

    New AI jailbreak methods exploit temporal and inscriptive vulnerabilities

    Researchers have developed new methods to bypass safety filters in AI models, targeting both large vision-language models (LVLMs) and text-to-image (T2I) models. One technique, TempJail, exploits temporal vulnerabilitie…

  10. RESEARCH · CL_216380 ·

    New research explores advanced jailbreak techniques and detection methods for LLMs and VLMs

    Researchers are developing advanced methods to test the safety and robustness of large language and vision-language models against jailbreaking attempts. New frameworks like SEAV focus on validating the correctness and …

  11. TOOL · CL_210440 ·

    New TTSD-FAR method enhances emotion recognition in LVLMs with missing data

    Researchers have developed a new method called TTSD-FAR for improving emotion recognition in large video-language models (LVLMs), particularly when some data modalities are missing during testing. This approach combines…

  12. RESEARCH · CL_198083 ·

    New LookBack method improves scoring for vision-language models

    Researchers have developed a new method called LookBack to better score the responses of Large Vision-Language Models (LVLMs). Existing methods, adapted from standard language models, struggle to evaluate how well an LV…

  13. TOOL · CL_196086 ·

    LVLMs struggle with temporal reasoning in image sequences, study finds

    A new research paper highlights a critical flaw in how current Large Vision-Language Models (LVLMs) evaluate temporal reasoning in image sequences. The study reveals that these models exhibit significant biases, such as…

  14. TOOL · CL_195987 ·

    SafeCap framework enhances LVLM safety via caption-mediated reinforcement learning

    Researchers have developed SafeCap, a new reinforcement learning framework designed to enhance the safety of Large Vision-Language Models (LVLMs). This method trains a policy model to generate safety-relevant image capt…

  15. TOOL · CL_201669 ·

    LVLMs struggle with temporal reasoning in image sequences due to placement bias

    A new paper highlights a critical flaw in how Large Vision-Language Models (LVLMs) evaluate temporal reasoning in image sequences. Current LVLMs, often used as judges, exhibit significant biases, favoring the placement …

  16. TOOL · CL_202788 ·

    SafeCap framework enhances LVLM safety using image captioning reinforcement learning

    Researchers have developed SafeCap, a novel reinforcement learning framework designed to enhance the safety of large vision-language models (LVLMs). SafeCap utilizes a learned self-captioning mechanism, where the model …

  17. TOOL · CL_191279 ·

    New Research Evaluates Confidence in Financial LVLMs

    A new arXiv paper explores confidence estimation for financial Vision-Language Models (LVLMs) used in chart and document understanding. The research highlights that while many models can rank correct answers above incor…

  18. RESEARCH · CL_172065 ·

    LVLMs struggle with visual illusions, new research reveals

    Researchers are investigating the limitations of Large Vision Language Models (LVLMs) in understanding visual illusions. One study proposes using visual illusions as a diagnostic tool to evaluate the joint perception an…

  19. RESEARCH · CL_167243 ·

    New methods tackle multimodal misinformation detection · 2 sources tracked

    Researchers have developed new methods for detecting multimodal misinformation, addressing challenges posed by both human-crafted and AI-generated deceptive content. One approach, Verification-Notebook Learning (VNL), u…

  20. RESEARCH · CL_165135 ·

    New research tackles LLM hallucinations across legal, multimodal, and general text generation

    Multiple research papers published on arXiv explore methods for detecting and mitigating hallucinations in large language models (LLMs). One study benchmarks legal hallucination detection, finding that while newer model…