Qwen2-Audio
PulseAugur coverage of Qwen2-Audio — every cluster mentioning Qwen2-Audio across labs, papers, and developer communities, ranked by signal.
3 day(s) with sentiment data
-
LLM-Augmented Alignment Improves Zero-Shot Respiratory Sound Classification
Researchers have developed a novel framework for zero-shot respiratory sound classification by aligning self-supervised respiratory encoders with clinical terminology. This method utilizes a medical LLM to generate stru…
-
New method audits generative audio LLMs, finds limited value in generative calls
Researchers have developed a new method to evaluate generative audio large language models (LLMs) by auditing their generative calls. This approach distinguishes between the value of acoustic evidence and the necessity …
-
New framework ARENA automates red-teaming for audio language models
Researchers have developed ARENA, a novel closed-loop framework designed for automated red-teaming of large audio-language models (LALMs). This system addresses the unique safety challenges posed by LALMs, which can exh…
-
New framework reveals cross-modal computation in multilingual speech-text models
Researchers have developed a new framework to analyze how multilingual speech-text models handle cross-modal language alignment. This framework, applied to models like SeamlessM4T and Qwen2-Audio, identifies language-se…
-
Audio LLM safety evaluated by speech delivery, not just text
Researchers have developed a new method called PJ-Break to evaluate the safety of audio-capable large language models (LLMs) by focusing on speech delivery rather than just transcript content. This method uses six speec…
-
New RL method boosts code-switched ASR data efficiency
Researchers have developed a novel reinforcement learning technique, RLVR, to improve the data efficiency of audio-language models for code-switched Automatic Speech Recognition (ASR). This method utilizes group relativ…
-
New CoAT Framework Enhances Large Audio Language Models with Continuous Thinking Space
Researchers have developed a new framework called Continuous Audio Thinking (CoAT) designed to enhance the capabilities of Large Audio Language Models (LALMs). CoAT equips these models with a continuous latent workspace…
-
New metric ALAS evaluates audio-language model alignment
Researchers have developed ALAS, an Automatic Latent Alignment Score, to evaluate how well audio language models align audio frames with text tokens. This model- and task-agnostic metric analyzes an LLM's hidden states,…
-
New methods probe and steer attention in audio AI models
Researchers have developed new methods to understand and manipulate the internal workings of large audio-language models. One technique, instruction-based vector steering, allows for the redirection of temporal attentio…
-
New method compresses audio tokens for language models
Researchers have developed a new method called Local Temporal Bipartite Merging (LTBM) to compress audio tokens in audio-language models. This training-free approach merges similar nearby audio tokens within a temporal …