PulseAugur
EN
LIVE 15:16:09
ENTITY Llava

Llava

PulseAugur coverage of Llava — every cluster mentioning Llava across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
5
21 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
4
16 over 90d
TIER MIX · 90D
TOPICS
SENTIMENT · 30D

3 day(s) with sentiment data

RECENT · PAGE 1/3 · 50 TOTAL
  1. TOOL · CL_257180 ·

    New ViD Framework Tackles Gender Bias in Vision-Language Models

    Researchers have introduced ViD, a novel framework designed to mitigate gender bias in large vision-language models (LVLMs). Unlike previous methods that require training-phase adjustments or post-hoc calibration, ViD a…

  2. TOOL · CL_254414 ·

    Unified Vision-Language Model Enhances PSMA PET/CT Analysis

    Researchers have developed a novel unified vision-language model designed to enhance the analysis of PSMA PET/CT scans for prostate cancer management. This model integrates report generation, visual question answering, …

  3. TOOL · CL_254250 ·

    MLLMs Mimic Human Perception of Bistable Images

    Researchers have investigated whether multimodal large language models (MLLMs) exhibit human-like reporting behavior when presented with bistable images, such as the classic duck-rabbit illusion. Using the LLaVA family …

  4. TOOL · CL_227256 ·

    Image augmentation techniques tested as generators for deep learning image retrieval systems

    This paper introduces a novel approach to testing deep learning-based image retrieval systems by utilizing image augmentation techniques as test generators. The research categorizes 50 augmentation methods and empirical…

  5. TOOL · CL_212130 ·

    New ClustRS Algorithm Boosts VLM Efficiency and Robustness

    Researchers have developed ClustRS, a novel two-part, training-free algorithm designed to enhance the efficiency and robustness of Visual-Language Models (VLMs). This method employs an attention-weighted clustering appr…

  6. TOOL · CL_203728 ·

    Developer runs multiple AI models locally via sequential loading

    A developer details a strategy for running multiple large AI models on a single local server with limited VRAM by employing a sequential loading approach. This method involves loading a model, using it for a specific ta…

  7. COMMENTARY · CL_200645 ·

    Stable Diffusion user seeks advice on image generation consistency

    A user on Reddit's r/StableDiffusion subreddit is seeking advice on how to achieve specific composition control and product consistency in image generation. They are comparing results from various models including ChatG…

  8. RESEARCH · CL_194006 ·

    New methods enhance open-vocabulary segmentation using multimodal pseudo-labels and test-time adaptation

    Researchers have developed new methods for open-vocabulary instance and panoptic segmentation, which aim to recognize objects beyond predefined categories without extensive manual annotation. One approach, detailed in a…

  9. TOOL · CL_201665 ·

    New Spanish-language cybersecurity vision-language model released

    Researchers have developed VectraYX-Vision-1B, a compact vision-language model designed for Spanish and Latin American cybersecurity imagery. This model, under 2 billion parameters, integrates a SigLIP-SO400M encoder wi…

  10. RESEARCH · CL_193699 ·

    New Spanish Cybersecurity Vision-Language Model Shows Promise Despite Grounding Issues

    Researchers have developed VectraYX-Vision-1B, a vision-language model designed for Spanish and Latin American cybersecurity imagery. This sub-2 billion parameter model integrates a SigLIP encoder with a Spanish securit…

  11. RESEARCH · CL_185181 ·

    New research explores multimodal alignment via optimal transport and latent denoising

    Two new research papers explore methods for improving multimodal alignment in large models. The first paper introduces Joint Kernel Entropic Gromov--Wasserstein Optimal Transport (JK-EGW) to align data from different mo…

  12. RESEARCH · CL_179099 ·

    New frameworks enhance VLM reasoning with visual tokens and self-diagnosis · 3 sources tracked

    Researchers have developed new frameworks to enhance the reasoning capabilities of Vision-Language Models (VLMs). One approach, Chain-of-Visual-Thought (COVT), uses continuous visual tokens to capture dense perceptual i…

  13. RESEARCH · CL_165240 ·

    New methods prune visual tokens for efficient MLLM inference · 4 sources tracked

    Researchers have developed several new methods to efficiently prune visual tokens for multimodal large language models (MLLMs), aiming to reduce inference costs and latency. The LAST framework uses the last query token'…

  14. TOOL · CL_162795 ·

    New framework boosts LVLM robustness against adversarial attacks

    Researchers have developed a dual adversarial fine-tuning framework to improve the robustness of Large Vision-Language Models (LVLMs) like LLaVA and GPT-4V against adversarial attacks. This new method enhances generaliz…

  15. TOOL · CL_150103 ·

    User seeks best practices for training custom Minecraft skin generator

    A user is seeking guidance on best practices for training a custom text-to-image and image-to-image model for generating Minecraft skins. They have curated a dataset of approximately 7,000 skins and are exploring variou…

  16. RESEARCH · CL_131401 ·

    VLMs enhance traffic sign assessment, outperforming manual methods

    Researchers have developed a new framework utilizing three fine-tuned Vision Language Models (VLMs) to comprehensively assess traffic sign conditions. This system integrates daytime visual performance, evaluating legibi…

  17. COMMENTARY · CL_126618 ·

    LocalLLaMA community seeks top open-weight VLMs for July 2026

    A Reddit discussion on the r/LocalLLaMA subreddit is seeking community input on the best locally runnable Vision Language Models (VLMs) as of July 2026. Participants are encouraged to share their preferred models, detai…

  18. TOOL · CL_119593 ·

    SMART framework optimizes speculative decoding for LLMs, boosting speed

    Researchers have developed SMART, a system-aware framework designed to optimize the efficiency of speculative decoding in large language models. This approach addresses the computational overhead that can lead to decrea…

  19. TOOL · CL_118151 ·

    New 'Neural Gate' method enhances LVLM privacy by editing neurons

    Researchers have developed a new method called Neural Gate to enhance the privacy of Large Vision-Language Models (LVLMs). This technique uses neuron-level model editing to identify and modify parameters associated with…

  20. TOOL · CL_115539 ·

    New BYORn Framework Defends LVLMs Against Backdoor Attacks

    Researchers have developed a novel defense framework called BYORn (Bootstrap Your Own Responses) to protect Large Vision-Language Models (LVLMs) from backdoor attacks during supervised fine-tuning (SFT). This method lev…