LLaVA-NeXT
PulseAugur coverage of LLaVA-NeXT — every cluster mentioning LLaVA-NeXT across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
Visual token pruning impacts MLLM calibration, research finds
A new research paper investigates the impact of visual token pruning on the calibration of multimodal large language models (MLLMs). The study, published on arXiv, reveals that the method used for pruning tokens signifi…
-
CRISP framework boosts LVLM efficiency by pruning visual tokens
Researchers have developed CRISP, a novel framework designed to improve the efficiency of Large Vision-Language Models (LVLMs) by pruning visual tokens before they are processed by the language model. This two-stage app…
-
Video2Reaction dataset maps video to audience emotions · 2 sources tracked
Researchers have introduced Video2Reaction, a new multimodal dataset and benchmark designed to predict audience emotional responses to video content. The dataset, comprising over 10,000 videos, utilizes a two-stage pipe…
-
New research tackles LLM reasoning, long-context, and tool integration
Multiple research papers explore advancements in large language model (LLM) reasoning capabilities, focusing on improving performance in long-horizon tasks and tool integration. Apple's research introduces LEAD, a metho…
-
ReasonCLIP-58M enhances CLIP models with visual commonsense reasoning
Researchers have introduced ReasonCLIP-58M, a new framework for continually pretraining CLIP-style models. This approach integrates large-scale reasoning supervision to enhance visually grounded commonsense inference an…
-
New TOPS method prunes visual tokens for efficient MLLM inference
Researchers have developed TOPS, a novel method for pruning visual tokens in multimodal large language models (MLLMs) to improve efficiency. Unlike previous approaches that relied on attention scores or token similarity…
-
New ALVTS method boosts LVLM efficiency with adaptive token selection
Researchers have introduced Adaptive Layer-wise Visual Token Selection (ALVTS), a new framework designed to improve the efficiency of Large Vision-Language Models (LVLMs). Unlike previous methods that permanently discar…
-
New AI Research Focuses on Model Efficiency via Quantization and Token Pruning
Researchers are developing new methods to improve the efficiency of AI models through quantization and token pruning. One approach, PeRQ, enhances post-training quantization by redistributing activation mass before rota…