Qwen2.5 Omni
PulseAugur coverage of Qwen2.5 Omni — every cluster mentioning Qwen2.5 Omni across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
llama.cpp bug causes non-deterministic results for M-RoPE embedding batches
A bug in the llama.cpp library causes incorrect results when processing embedding batches for M-RoPE models like Qwen3.5 and Qwen2.5-VL. The issue stems from a heap buffer overflow where the library reads past the alloc…
-
Kraken LLM advances speech-to-speech translation quality
Researchers have developed Kraken, a novel speech-to-speech translation model that leverages LLMs and low-bitrate vector quantization for improved quality and non-linguistic information preservation. The model, built up…
-
New research explores LLM efficiency, safety, and multilingual capabilities
Researchers are exploring various methods to enhance the efficiency and capabilities of large language models (LLMs). Apple Inc. has published research on improving multilingual speech models by enhancing language discr…
-
AI models rely on prosody stereotypes for sarcasm detection, study finds
A new research paper investigates how multimodal large language models (MLLMs) process speech and text, specifically focusing on their detection of sarcasm. The study found that models like Qwen2.5 Omni, Qwen3-Omni, and…
-
New framework enhances AI spoken turn-taking with style awareness
Researchers have developed a generalized style-aware full-duplex framework to improve the timing and quality of spoken interactions in AI systems. The framework includes LPS-TC, a lightweight controller for proactive tu…
-
New method audits generative audio LLMs, finds limited value in generative calls
Researchers have developed a new method to evaluate generative audio large language models (LLMs) by auditing their generative calls. This approach distinguishes between the value of acoustic evidence and the necessity …
-
Multimodal LLMs Rely on Sarcasm Heuristics, Not True Prosody
Researchers investigated how multimodal large language models (MLLMs) process speech and text, specifically focusing on sarcasm detection. Their experiments with Qwen2.5-Omni and Qwen3-Omni revealed that adding audio in…
-
AI models' emotion neurons identified and validated across languages
Researchers have conducted the first neuron-level interpretability studies on large audio-language models (LALMs) to understand how they encode emotion across different languages. The studies identified "Multilingual Em…
-
New methods slash OmniLLM token costs, boosting efficiency and accuracy · 9 sources tracked
Researchers have developed several novel methods for compressing token sequences in omnimodal large language models (OmniLLMs) to reduce memory and inference costs. These approaches, including OmniDelta, OmniScope, Prog…
-
New IAAN method boosts LALM acoustic perception by targeting encoder neurons
Researchers have developed a novel method called IAAN (Identifying and Amplifying Acoustic Neurons) to enhance the acoustic perception capabilities of large audio-language models (LALMs) without requiring retraining. Th…
-
New methods tackle OmniLLM token compression for efficiency
Two new research papers propose methods to compress token sequences in omnimodal large language models (OmniLLMs) to reduce inference costs. The first paper, DASH, uses audio cues to dynamically segment sequences and a …
-
FlexiSLM introduces dynamic frame rates for spoken language models
Researchers have developed FlexiSLM, a novel spoken language model that dynamically adjusts its frame rate for speech input and output. Unlike existing models that use fixed frame rates, FlexiSLM can adapt to the varyin…
-
Researchers map audio-visual information flow in multimodal LLMs
Researchers have investigated the internal information flow within multimodal large language models (MLLMs) that process both audio and visual data. Their study, focusing on Audio-Visual Large Language Models (AVLLMs), …
-
New research finds modality alignment transfers AI audio attacks
A new research paper introduces the "Alignment Curse," a principle demonstrating how improved text-audio modality alignment in omni-models can inadvertently transfer safety vulnerabilities from text to audio. Researcher…
-
New LLM tools automate speech and transcript error annotation
Researchers have developed two new methods for automatically annotating errors in transcribed speech. One approach, Speech Translation Error Labelling (STEL), uses existing text-only and multimodal LLMs to identify erro…
-
Raon-Speech launches 9B parameter model for speech understanding and generation
Researchers have introduced Raon-Speech, a 9-billion parameter speech language model capable of understanding, answering, and generating speech in English and Korean. This model, trained on over 1.38 million hours of cu…
-
SEATS method slashes LLM compute by pruning audio-visual tokens
Researchers have developed SEATS, a new method to make omni-modal large language models (om-LLMs) more efficient. SEATS prunes redundant audio-visual tokens throughout the model's layers, adapting the token selection pr…
-
AffectVerse model predicts future emotions using temporal imagination
Researchers have introduced AffectVerse, a new multimodal model designed for affective computing that integrates temporal prediction into its reasoning process. Unlike previous models that treated emotion recognition st…
-
Omni-Encoder unifies vision and audio processing for human-like motion perception
Researchers have developed Omni-Encoder, a novel Transformer backbone that unifies visual and audio signals for more holistic perception. Unlike previous models that process modalities separately and at different rates,…
-
New framework reveals audio hallucinations in egocentric video models
Researchers have developed a new framework to evaluate audio hallucinations in egocentric videos, where models infer sounds from visual cues that are not actually heard. Their study found that advanced audio-visual lang…