LLaVA-onevision
PulseAugur coverage of LLaVA-onevision — every cluster mentioning LLaVA-onevision across labs, papers, and developer communities, ranked by signal.
-
GeoTrace framework compresses video tokens for efficient Video LLMs
Researchers have introduced GeoTrace, a novel framework designed to enhance the efficiency of Video Large Language Models (Video LLMs) by compressing visual tokens. This training-free method decomposes video evidence in…
-
New method tracks multimodal LLM attention token-by-token
Researchers have developed a new method called "One Token at a Time" (OTaT) to analyze how multimodal large language models (MLLMs) utilize visual and textual information during response generation. This technique track…
-
New framework boosts high-resolution image perception in LLMs
Researchers have introduced Hierarchical Entity Exploration (HEE), a novel framework designed to enhance high-resolution image perception in multimodal large language models (MLLMs). Unlike existing methods that require…
-
New frameworks enhance multimodal AI by preserving knowledge and improving generation
Researchers are developing new frameworks to enhance multimodal AI models. Rosetta introduces a composable pretraining approach that preserves core knowledge while adding new modalities non-destructively, using Momentum…
-
New methods enhance streaming video understanding with efficient memory and re-watch capabilities · 6 sources tracked
Researchers have developed new methods to improve streaming video understanding (SVU) under strict computational and memory constraints. ProtoKV, a novel memory system, aggregates older video content into a summary stat…
-
New research advances vector quantization for AI models
Several recent research papers explore advancements in vector quantization techniques for AI models. ArcVQ-VAE introduces a spherical angular-margin prior to improve latent representation diversity and codebook utilizat…