Large Vision Language Models
PulseAugur coverage of Large Vision Language Models — every cluster mentioning Large Vision Language Models across labs, papers, and developer communities, ranked by signal.
6 day(s) with sentiment data
-
Visual framing impacts LVLM descriptions, study finds
A new paper published on arXiv explores how the visual input and its framing can influence the attribute-based descriptions generated by large vision-language models (LVLMs). The research indicates that even when a text…
-
New ViD Framework Tackles Gender Bias in Vision-Language Models
Researchers have introduced ViD, a novel framework designed to mitigate gender bias in large vision-language models (LVLMs). Unlike previous methods that require training-phase adjustments or post-hoc calibration, ViD a…
-
New framework ExpArt-KG enhances LVLMs for artwork description
Researchers have developed ExpArt-KG, a framework designed to enhance the descriptive capabilities of Large Vision-Language Models (LVLMs) for artwork. This method integrates knowledge graphs with retrieval-augmented ge…
-
New framework and benchmark improve diagram-to-graph topology extraction
Researchers have introduced TopoAgent, a novel framework designed to improve the extraction of graph topologies from structural diagrams using large vision-language models. This framework is accompanied by TopoBench-180…
-
New dataset Synth-JDoc boosts LVLM Japanese OCR performance
Researchers have developed Synth-JDoc, a novel dataset designed to improve the optical character recognition (OCR) capabilities of Large Vision Language Models (LVLMs), particularly for Japanese text. This dataset synth…
-
New benchmarks and methods tackle LLM hallucinations across modalities and domains
Researchers are developing new methods and benchmarks to detect and mitigate hallucinations in large language models (LLMs) across various modalities and domains. OmniHallu offers a unified framework for detecting hallu…
-
New research tackles AI hallucinations in video and language models
Researchers are developing new methods to combat hallucinations in AI models, particularly in video-language and large language models. One approach, CounterVid, uses counterfactual video generation to create synthetic …
-
New benchmark and method tackle cross-lingual VQA degradation in medical AI
Researchers have developed a new benchmark and a method to address cross-lingual degradation in multilingual medical visual question answering (VQA). The benchmark, covering eight languages and four scenarios, reveals t…
-
New SAGE framework enhances ancient document understanding with multi-agent inference
Researchers have developed SAGE, a novel multi-agent framework designed to improve the understanding of Chinese ancient documents. Unlike current Large Vision-Language Models (LVLMs) that generate answers directly, SAGE…
-
Re$^3$Cap uses retrieval-guided reinforcement learning to improve image captioning
Researchers have developed Re$^3$Cap, a novel method for enhancing image captioning by leveraging reinforcement learning guided by multi-modal retrieval. This approach aims to overcome the limitations of standard reinfo…
-
New AI jailbreak methods exploit temporal and inscriptive vulnerabilities
Researchers have developed new methods to bypass safety filters in AI models, targeting both large vision-language models (LVLMs) and text-to-image (T2I) models. One technique, TempJail, exploits temporal vulnerabilitie…
-
New research explores advanced jailbreak techniques and detection methods for LLMs and VLMs
Researchers are developing advanced methods to test the safety and robustness of large language and vision-language models against jailbreaking attempts. New frameworks like SEAV focus on validating the correctness and …
-
TAMP-Nav framework enhances embodied navigation for Large Vision-Language Models
Researchers have introduced TAMP-Nav, a new framework designed to enhance embodied navigation for Large Vision-Language Models (VLMs). This approach reformulates navigation tasks into 2D visual prompting, allowing VLMs …
-
CityRiSE framework enhances LVLMs for urban socio-economic prediction
Researchers have developed CityRiSE, a new framework that uses reinforcement learning to improve the ability of Large Vision-Language Models (LVLMs) to predict urban socio-economic status. This approach guides LVLMs to …
-
New SGU framework holistically evaluates unified multimodal AI models
Researchers have introduced Self-Generative-Understanding (SGU), a new evaluation framework designed to holistically assess unified multimodal models (UMMs). Unlike existing methods that evaluate generative and discrimi…
-
New framework evaluates unified multimodal AI models holistically
Researchers have introduced Self-Generative-Understanding (SGU), a new framework for evaluating unified multimodal models (UMMs). Current methods often assess visual generation and understanding separately, failing to c…
-
New LookBack method improves scoring for vision-language models
Researchers have developed a new method called LookBack to better score the responses of Large Vision-Language Models (LVLMs). Existing methods, adapted from standard language models, struggle to evaluate how well an LV…
-
LVLMs struggle with temporal reasoning in image sequences, study finds
A new research paper highlights a critical flaw in how current Large Vision-Language Models (LVLMs) evaluate temporal reasoning in image sequences. The study reveals that these models exhibit significant biases, such as…
-
LVLMs struggle with temporal reasoning in image sequences due to placement bias
A new paper highlights a critical flaw in how Large Vision-Language Models (LVLMs) evaluate temporal reasoning in image sequences. Current LVLMs, often used as judges, exhibit significant biases, favoring the placement …
-
New frameworks tackle hallucination in multimodal AI models · 3 sources tracked
Researchers have developed new frameworks to combat hallucinations in multimodal large language models (MLLMs). UniHall introduces a fine-grained dataset and a self-adaptive fuzzing framework (SAMF) to stress-test MLLMs…