Llava
PulseAugur coverage of Llava — every cluster mentioning Llava across labs, papers, and developer communities, ranked by signal.
3 day(s) with sentiment data
-
New ViD Framework Tackles Gender Bias in Vision-Language Models
Researchers have introduced ViD, a novel framework designed to mitigate gender bias in large vision-language models (LVLMs). Unlike previous methods that require training-phase adjustments or post-hoc calibration, ViD a…
-
Unified Vision-Language Model Enhances PSMA PET/CT Analysis
Researchers have developed a novel unified vision-language model designed to enhance the analysis of PSMA PET/CT scans for prostate cancer management. This model integrates report generation, visual question answering, …
-
MLLMs Mimic Human Perception of Bistable Images
Researchers have investigated whether multimodal large language models (MLLMs) exhibit human-like reporting behavior when presented with bistable images, such as the classic duck-rabbit illusion. Using the LLaVA family …
-
Image augmentation techniques tested as generators for deep learning image retrieval systems
This paper introduces a novel approach to testing deep learning-based image retrieval systems by utilizing image augmentation techniques as test generators. The research categorizes 50 augmentation methods and empirical…
-
New ClustRS Algorithm Boosts VLM Efficiency and Robustness
Researchers have developed ClustRS, a novel two-part, training-free algorithm designed to enhance the efficiency and robustness of Visual-Language Models (VLMs). This method employs an attention-weighted clustering appr…
-
Developer runs multiple AI models locally via sequential loading
A developer details a strategy for running multiple large AI models on a single local server with limited VRAM by employing a sequential loading approach. This method involves loading a model, using it for a specific ta…
-
Stable Diffusion user seeks advice on image generation consistency
A user on Reddit's r/StableDiffusion subreddit is seeking advice on how to achieve specific composition control and product consistency in image generation. They are comparing results from various models including ChatG…
-
New methods enhance open-vocabulary segmentation using multimodal pseudo-labels and test-time adaptation
Researchers have developed new methods for open-vocabulary instance and panoptic segmentation, which aim to recognize objects beyond predefined categories without extensive manual annotation. One approach, detailed in a…
-
New Spanish-language cybersecurity vision-language model released
Researchers have developed VectraYX-Vision-1B, a compact vision-language model designed for Spanish and Latin American cybersecurity imagery. This model, under 2 billion parameters, integrates a SigLIP-SO400M encoder wi…
-
New Spanish Cybersecurity Vision-Language Model Shows Promise Despite Grounding Issues
Researchers have developed VectraYX-Vision-1B, a vision-language model designed for Spanish and Latin American cybersecurity imagery. This sub-2 billion parameter model integrates a SigLIP encoder with a Spanish securit…
-
New research explores multimodal alignment via optimal transport and latent denoising
Two new research papers explore methods for improving multimodal alignment in large models. The first paper introduces Joint Kernel Entropic Gromov--Wasserstein Optimal Transport (JK-EGW) to align data from different mo…
-
New frameworks enhance VLM reasoning with visual tokens and self-diagnosis · 3 sources tracked
Researchers have developed new frameworks to enhance the reasoning capabilities of Vision-Language Models (VLMs). One approach, Chain-of-Visual-Thought (COVT), uses continuous visual tokens to capture dense perceptual i…
-
New methods prune visual tokens for efficient MLLM inference · 4 sources tracked
Researchers have developed several new methods to efficiently prune visual tokens for multimodal large language models (MLLMs), aiming to reduce inference costs and latency. The LAST framework uses the last query token'…
-
New framework boosts LVLM robustness against adversarial attacks
Researchers have developed a dual adversarial fine-tuning framework to improve the robustness of Large Vision-Language Models (LVLMs) like LLaVA and GPT-4V against adversarial attacks. This new method enhances generaliz…
-
User seeks best practices for training custom Minecraft skin generator
A user is seeking guidance on best practices for training a custom text-to-image and image-to-image model for generating Minecraft skins. They have curated a dataset of approximately 7,000 skins and are exploring variou…
-
VLMs enhance traffic sign assessment, outperforming manual methods
Researchers have developed a new framework utilizing three fine-tuned Vision Language Models (VLMs) to comprehensively assess traffic sign conditions. This system integrates daytime visual performance, evaluating legibi…
-
LocalLLaMA community seeks top open-weight VLMs for July 2026
A Reddit discussion on the r/LocalLLaMA subreddit is seeking community input on the best locally runnable Vision Language Models (VLMs) as of July 2026. Participants are encouraged to share their preferred models, detai…
-
SMART framework optimizes speculative decoding for LLMs, boosting speed
Researchers have developed SMART, a system-aware framework designed to optimize the efficiency of speculative decoding in large language models. This approach addresses the computational overhead that can lead to decrea…
-
New 'Neural Gate' method enhances LVLM privacy by editing neurons
Researchers have developed a new method called Neural Gate to enhance the privacy of Large Vision-Language Models (LVLMs). This technique uses neuron-level model editing to identify and modify parameters associated with…
-
New BYORn Framework Defends LVLMs Against Backdoor Attacks
Researchers have developed a novel defense framework called BYORn (Bootstrap Your Own Responses) to protect Large Vision-Language Models (LVLMs) from backdoor attacks during supervised fine-tuning (SFT). This method lev…