LLaVA-NeXT-7B
PulseAugur coverage of LLaVA-NeXT-7B — every cluster mentioning LLaVA-NeXT-7B across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
New methods accelerate Vision-Language Model inference by optimizing token processing · 2 sources tracked
Two new research papers propose methods to accelerate the inference of Vision-Language Models (VLMs) by reducing computational overhead. StackTok focuses on adaptive visual token selection, prioritizing query relevance …
-
New dataset and framework tackle multi-modal LLM safety in conversations
Researchers have introduced MINT-Safe, a new dataset designed to address safety concerns in multi-modal large language models (MLLMs) during extended conversational interactions. This dataset, comprising 11,270 multi-im…
-
New benchmark reveals safety reasoning gap in vision-language models
A new benchmark called SafeGesture has been developed to evaluate the fine-grained hand gesture understanding of vision-language models (VLMs) in safety-critical scenarios. The benchmark pairs six gestures with eight op…
-
New methods emerge for efficient visual token pruning in AI models · 6 sources tracked
Researchers are developing new methods to optimize Vision Transformers (ViTs) and Multimodal Large Language Models (MLLMs) by pruning visual tokens, which are computationally expensive. Several papers propose novel tech…
-
RUTA method drastically cuts visual tokens for LLMs while preserving performance
Researchers have developed RUTA, a novel method for reducing the number of visual tokens processed by large language models. RUTA learns to select and allocate tokens based on query-specific relevance and a rate-utility…
-
AnchorPrune framework enhances vision-language model efficiency by pruning tokens
Researchers have developed AnchorPrune, a novel framework designed to optimize the efficiency of large vision-language models by pruning redundant visual tokens. This training-free method constructs a relevance anchor a…
-
Perceptual Flow Network and VGR enhance visual reasoning in LLMs
Researchers have developed a Perceptual Flow Network (PFlowNet) to improve visual reasoning in Large-Vision Language Models (LVLMs). PFlowNet decouples perception from reasoning and uses variational reinforcement learni…