LLaVA-1.5
PulseAugur coverage of LLaVA-1.5 — every cluster mentioning LLaVA-1.5 across labs, papers, and developer communities, ranked by signal.
4 day(s) with sentiment data
-
New Wiener Filtering Technique Reduces Hallucinations in Vision-Language Models
Researchers have developed a novel technique called Wiener Representation Filtering to reduce hallucinations in vision-language models (VLMs). This training-free method operates post-hoc by editing the representation sp…
-
New AI Frameworks Tackle Visual Token Pruning in Multimodal LLMs
Researchers are developing new methods to optimize multimodal large language models (MLLMs) by pruning visual tokens, which are computationally expensive. One approach, MAP, predicts the importance of visual tokens by l…
-
New MoP framework compresses LLMs, boosting efficiency and accuracy
Researchers have developed a new iterative framework called Mixture of Pruners (MoP) designed to compress Large Language Models (LLMs) by reducing their parameter count and accelerating inference. MoP unifies depth and …
-
Visual token pruning impacts MLLM calibration, research finds
A new research paper investigates the impact of visual token pruning on the calibration of multimodal large language models (MLLMs). The study, published on arXiv, reveals that the method used for pruning tokens signifi…
-
CRISP framework boosts LVLM efficiency by pruning visual tokens
Researchers have developed CRISP, a novel framework designed to improve the efficiency of Large Vision-Language Models (LVLMs) by pruning visual tokens before they are processed by the language model. This two-stage app…
-
University of Michigan unveils NeuroVFM for neuroimaging analysis
Researchers at the University of Michigan have developed NeuroVFM, a novel foundation model for neuroimaging. Trained using the Vol-JEPA approach on over 5.24 million clinical MRI and CT scans, NeuroVFM learns from uncu…
-
New framework combats catastrophic forgetting in MLLMs
Researchers have introduced Curvature-Guided Mixing (CGM), a new framework designed to improve the adaptation of Multimodal Large Language Models (MLLMs). This method addresses the issue of catastrophic forgetting, wher…
-
New ALVTS method boosts LVLM efficiency with adaptive token selection
Researchers have introduced Adaptive Layer-wise Visual Token Selection (ALVTS), a new framework designed to improve the efficiency of Large Vision-Language Models (LVLMs). Unlike previous methods that permanently discar…
-
Vision-language model predicts coastlines as polylines
Researchers have developed CoastlineVLM-7B, a vision-language model designed to directly predict coastlines as polylines rather than segmentation masks. This approach, built on the GeoChat-7B/LLaVA-1.5 architecture, foc…
-
VG-CoT: Towards Trustworthy Visual Reasoning via Grounded Chain-of-Thought
Researchers have introduced VG-CoT, a new dataset designed to improve the trustworthiness of Large Vision-Language Models (LVLMs). This dataset automatically links reasoning steps to specific visual evidence within imag…