Qwen2.5 Omni
PulseAugur coverage of Qwen2.5 Omni — every cluster mentioning Qwen2.5 Omni across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
AI models' emotion neurons identified and validated across languages
Researchers have conducted the first neuron-level interpretability studies on large audio-language models (LALMs) to understand how they encode emotion across different languages. The studies identified "Multilingual Em…
-
New methods slash OmniLLM token costs, boosting efficiency and accuracy · 9 sources tracked
Researchers have developed several novel methods for compressing token sequences in omnimodal large language models (OmniLLMs) to reduce memory and inference costs. These approaches, including OmniDelta, OmniScope, Prog…
-
New IAAN method boosts LALM acoustic perception by targeting encoder neurons
Researchers have developed a novel method called IAAN (Identifying and Amplifying Acoustic Neurons) to enhance the acoustic perception capabilities of large audio-language models (LALMs) without requiring retraining. Th…
-
New methods tackle OmniLLM token compression for efficiency
Two new research papers propose methods to compress token sequences in omnimodal large language models (OmniLLMs) to reduce inference costs. The first paper, DASH, uses audio cues to dynamically segment sequences and a …
-
FlexiSLM introduces dynamic frame rates for spoken language models
Researchers have developed FlexiSLM, a novel spoken language model that dynamically adjusts its frame rate for speech input and output. Unlike existing models that use fixed frame rates, FlexiSLM can adapt to the varyin…
-
Researchers map audio-visual information flow in multimodal LLMs
Researchers have investigated the internal information flow within multimodal large language models (MLLMs) that process both audio and visual data. Their study, focusing on Audio-Visual Large Language Models (AVLLMs), …
-
New research finds modality alignment transfers AI audio attacks
A new research paper introduces the "Alignment Curse," a principle demonstrating how improved text-audio modality alignment in omni-models can inadvertently transfer safety vulnerabilities from text to audio. Researcher…
-
New LLM tools automate speech and transcript error annotation
Researchers have developed two new methods for automatically annotating errors in transcribed speech. One approach, Speech Translation Error Labelling (STEL), uses existing text-only and multimodal LLMs to identify erro…
-
Raon-Speech launches 9B parameter model for speech understanding and generation
Researchers have introduced Raon-Speech, a 9-billion parameter speech language model capable of understanding, answering, and generating speech in English and Korean. This model, trained on over 1.38 million hours of cu…
-
SEATS method slashes LLM compute by pruning audio-visual tokens
Researchers have developed SEATS, a new method to make omni-modal large language models (om-LLMs) more efficient. SEATS prunes redundant audio-visual tokens throughout the model's layers, adapting the token selection pr…
-
AffectVerse model predicts future emotions using temporal imagination
Researchers have introduced AffectVerse, a new multimodal model designed for affective computing that integrates temporal prediction into its reasoning process. Unlike previous models that treated emotion recognition st…
-
Omni-Encoder unifies vision and audio processing for human-like motion perception
Researchers have developed Omni-Encoder, a novel Transformer backbone that unifies visual and audio signals for more holistic perception. Unlike previous models that process modalities separately and at different rates,…
-
New framework reveals audio hallucinations in egocentric video models
Researchers have developed a new framework to evaluate audio hallucinations in egocentric videos, where models infer sounds from visual cues that are not actually heard. Their study found that advanced audio-visual lang…