Speech LLMs
PulseAugur coverage of Speech LLMs — every cluster mentioning Speech LLMs across labs, papers, and developer communities, ranked by signal.
3 day(s) with sentiment data
-
New EAVA Method Enhances Speech-LLM Adaptation for ASR
Researchers have developed a new method called Encoder Awakening via Adapters (EAVA) to improve the domain-adaptive fine-tuning of Speech Large Language Models (Speech-LLMs) for Automatic Speech Recognition (ASR). This …
-
New \"$\alpha$-split\" method enhances privacy for speech LLMs in federated learning
Researchers have developed a new method called \"$\alpha$-split\" to improve differential privacy in federated learning for multilingual speech large language models (speech-LLMs). Standard per-layer differential privac…
-
X-AuT framework progressively compresses speech LLM audio encoders
Researchers have developed X-AuT, a novel framework designed to progressively compress audio-encoder layers in speech large language models. This method aims to reduce inference costs without significantly sacrificing a…
-
New WnW KV Cache Method Optimizes LLMs for Long-Form Speech
Researchers have developed a novel method called Waxing-and-Waning KV cache (WnW) to optimize memory usage in large language models designed for long-form speech processing. This technique categorizes KV cache heads int…
-
Speech LLMs for Low-Resource Languages: New Research Explores Data Needs and Pretraining
A new research paper explores the effectiveness of Speech Large Language Models (LLMs) for Automatic Speech Recognition (ASR) in low-resource languages. The study, utilizing the SLAM-ASR framework, assesses the data vol…
-
New gradient-based method aligns speech-to-text across all ASR models
Researchers have developed a novel gradient-based method for aligning speech-to-text, applicable to any differentiable automatic speech recognition (ASR) model. This technique derives word timings from the gradient of t…
-
Speech LLMs enhanced by translation-focused encoder pre-training
Researchers have explored a novel approach to enhance Speech LLMs by integrating translation objectives into the pre-training of speech encoders. This method addresses the structural misalignment between language-specif…
-
Speech LLMs enhanced by translation-based encoder pre-training
A new research paper proposes using speech translation to bridge the gap between speech encoders and large language models (LLMs) in Speech LLMs. The paper argues that current architectures have a structural misalignmen…
-
New framework SURE standardizes speech AI model evaluation
Researchers have introduced SURE, a unified framework designed to standardize and improve the reproducibility of speech understanding model evaluations. This framework addresses the challenge of comparing different spee…