Large Audio-Language Models
PulseAugur coverage of Large Audio-Language Models — every cluster mentioning Large Audio-Language Models across labs, papers, and developer communities, ranked by signal.
5 day(s) with sentiment data
-
New framework measures fairness in audio language models
Researchers have developed a new framework to evaluate fairness in Large Audio Language Models (LALMs). This semantic-aware mixed-effects regression approach addresses challenges in spoken-input settings by accounting f…
-
AI models' emotion neurons identified and validated across languages
Researchers have conducted the first neuron-level interpretability studies on large audio-language models (LALMs) to understand how they encode emotion across different languages. The studies identified "Multilingual Em…
-
New research tackles LLM and LALM safety risks with probabilistic and low-frequency input analysis · 2 sources tracked
Researchers have introduced ProbGuard, a novel probabilistic approach to enhance Large Language Model (LLM) safety by leveraging early output distributional signals. This method aims to detect and mitigate unsafe genera…
-
HyPASE framework uses hyperbolic geometry for efficient LALM fine-tuning
Researchers have developed HyPASE, a novel framework that utilizes hyperbolic geometry for parameter-efficient fine-tuning of Large Audio-Language Models (LALMs) for Speech Emotion Recognition (SER). Unlike traditional …
-
ESCUCHA benchmark launched for Spanish speech understanding in LALMs
Researchers have introduced ESCUCHA, a new benchmark designed to evaluate large audio language models (LALMs) specifically for the Spanish language. This benchmark addresses a gap in evaluating LALMs under realistic, he…
-
Large Audio Language Models Show Promise for Voice Authentication Systems
Researchers have explored the application of Large Audio Language Models (LALMs) for Spoofing-Aware Speaker Verification (SASV), a critical area for voice authentication systems facing advanced text-to-speech and voice …
-
New research tackles audio-language model limitations in instruction following and evaluation
Researchers are developing new methods to improve the capabilities of large audio-language models (LALMs). One approach focuses on using audio-aware LLMs to provide fine-grained feedback for better instruction following…
-
New research tackles audio deepfake detection with advanced AI techniques
Researchers are developing advanced methods for detecting audio deepfakes, focusing on improving generalization and interpretability. One approach, SONAR, uses a frequency-guided framework to exploit high-frequency arti…
-
New pipeline Auto-AEG boosts audio event localization for LALMs
Researchers have developed Auto-AEG, a scalable pipeline designed to construct supervision data for open-vocabulary audio event grounding. This task aims to precisely locate sound events described by natural language qu…
-
New VIBE framework reveals systematic bias in Large Audio-Language Models
A new framework called VIBE has been developed to evaluate biases in Large Audio-Language Models (LALMs) using real-world speech and open-ended tasks. Unlike previous methods that relied on synthetic speech or multiple-…
-
New method improves audio-language model accuracy with adaptive transformations
Researchers have developed a new method called Adaptive Perturbation Selection (APS) to improve the accuracy of large audio-language models (LALMs). Existing contrastive decoding techniques often use blunt methods like …
-
ALM2Vec framework uses large audio-language models for universal audio retrieval
Researchers have introduced ALM2Vec, a novel framework designed to create universal audio embeddings by leveraging large audio-language models (LALMs). Unlike previous methods focused on audio-caption matching, ALM2Vec …
-
New benchmark evaluates audio LLMs' context-aware scene understanding
Researchers have introduced a new benchmark called CASU (Context-Aware Auditory Scene Understanding) to evaluate Large Audio Language Models (LALMs). Existing benchmarks often assess audio layers like speech or sound in…
-
New benchmark reveals LALM judges lag human paralinguistic evaluation
Researchers have developed ParaPairAudioBench, a new benchmark designed to evaluate Large Audio-Language Models (LALMs) on their ability to distinguish subtle paralinguistic features in speech. The benchmark includes 5,…
-
New CoAT Framework Enhances Large Audio Language Models with Continuous Thinking Space
Researchers have developed a new framework called Continuous Audio Thinking (CoAT) designed to enhance the capabilities of Large Audio Language Models (LALMs). CoAT equips these models with a continuous latent workspace…
-
New benchmarks tackle privacy risks in large language models
Researchers have developed new methods to evaluate membership inference attacks (MIAs) against large language models (LLMs), particularly focusing on audio and text modalities. The first study introduces a systematic ev…
-
New AudioDER Dataset Boosts LALM Reasoning Capabilities
Researchers have introduced AudioDER, a new dataset designed to enhance the reasoning capabilities of Large Audio-Language Models (LALMs). The dataset addresses the issue of redundancy in existing audio-language dataset…
-
SpectCount uses synthetic audio to boost large audio language models
Researchers have developed SpectCount, a novel method for improving large audio language models (LALMs) by using synthetic audio signals. This approach addresses the scarcity of high-quality annotated audio data by gene…
-
New adapter adds test-time memory to audio LLMs for better emotion recognition
Researchers have developed a novel method called Titans-as-a-Layer (MAL) to enhance conversational speech emotion recognition. This plug-and-play adapter integrates test-time neural memory into large audio language mode…
-
New GlobeAudio benchmark tests AI audio models on naturalistic language
Researchers have introduced GlobeAudio, a new benchmark designed to evaluate Large Audio-Language Models (LALMs) in more realistic, naturalistic settings. The benchmark features 5,637 multiple-choice questions in six di…