Qwen2.5-Omni-3B
PulseAugur coverage of Qwen2.5-Omni-3B — every cluster mentioning Qwen2.5-Omni-3B across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
New Triage method optimizes audio token processing for LALMs
Researchers have developed a novel method called Triage for optimizing audio token processing in large audio language models (LALMs). Triage predicts the attention audio tokens will receive within the language model bef…
-
New WnW KV Cache Method Optimizes LLMs for Long-Form Speech
Researchers have developed a novel method called Waxing-and-Waning KV cache (WnW) to optimize memory usage in large language models designed for long-form speech processing. This technique categorizes KV cache heads int…
-
New A-PACK framework slashes omni-LLM costs with deferred audio pruning
Researchers have introduced A-PACK, a novel two-stage framework designed to reduce the computational costs associated with omni-modal Large Language Models (LLMs). This method defers audio pruning until query-conditione…
-
Gemma 4 E2B leads industrial edge AI model tests over faster rivals
A recent test of five small multimodal models on a Jetson device for an industrial edge AI runtime found that Gemma 4 E2B remained the baseline despite not being the fastest. While SmolVLM2 was the quickest, its outputs…
-
New frameworks and benchmarks advance Video-LLM efficiency and understanding
Researchers have introduced EarlyTom, a novel framework designed to enhance the efficiency of video large language models (Video-LLMs) by compressing visual tokens early in the vision encoder. This approach significantl…
-
HeadRouter prunes audio tokens in LLMs by routing attention heads
Researchers have introduced HeadRouter, a novel method for compressing large audio language models by dynamically pruning audio tokens. Unlike previous approaches that assume uniform head importance, HeadRouter recogniz…