speech recognition
PulseAugur coverage of speech recognition — every cluster mentioning speech recognition across labs, papers, and developer communities, ranked by signal.
16 day(s) with sentiment data
-
New EAVA Method Enhances Speech-LLM Adaptation for ASR
Researchers have developed a new method called Encoder Awakening via Adapters (EAVA) to improve the domain-adaptive fine-tuning of Speech Large Language Models (Speech-LLMs) for Automatic Speech Recognition (ASR). This …
-
New ASR method improves accuracy and latency
Researchers have developed a new method for streaming automatic speech recognition (ASR) that improves both transcription accuracy and latency. The approach, called AWED, uses a word-level emission-delay metric and a no…
-
New benchmark RoleBreak tests long-horizon role-playing in spoken dialogue systems
Researchers have introduced RoleBreak, a new benchmark designed to evaluate the long-horizon role-playing capabilities of spoken dialogue systems. The benchmark includes over 300 roles and thousands of human-verified di…
-
Open-source tools enable private, self-hosted voice AI assistants
An engineer named Ravi Roy highlights the growing trend of building private, open-source voice AI assistants, moving away from proprietary cloud services. This approach offers enhanced data privacy, reduced costs, and g…
-
New DUPAR framework enhances voice assistant retrieval speed and accuracy
Researchers have developed DUPAR, a novel conversational retrieval framework designed to improve the speed and accuracy of voice assistants. This system utilizes a dual-path approach: a fast path with an adapted audio e…
-
Machine unlearning techniques reduce privacy risks in audio-language models
Researchers have developed and evaluated several machine unlearning strategies for Large Audio-Language Models (LALMs) used in Speech Question Answering. These methods, including gradient ascent, task arithmetic, and al…
-
LLM merging enhances ASR performance without added computational cost
Researchers have developed a novel method for integrating large language models (LLMs) into automatic speech recognition (ASR) systems without increasing computational costs. This approach involves merging the LLMs dire…
-
Quranic text mapping and recitation validator released
Researchers have developed a new method for mapping Quranic text between its Uthmani and Standard Arabic orthographic forms, addressing discrepancies caused by the Unicode character U+0670. This work includes a 2,290-pa…
-
ASLP team develops end-to-end multimodal system for clinical SOAP note generation
Researchers from the ASLP team have developed a novel end-to-end multimodal system for generating structured SOAP notes directly from long-form clinical audio. This system, designed for the BeTraC 2026 challenge, bypass…
-
VideoXAgent tackles long video understanding with online agent harness
Researchers have developed VideoXAgent, an online harness designed for understanding long videos. This system plans tasks, uses specialized tools like VLMs, OCR, and ASR on demand, and aggregates evidence to answer quer…
-
New AI method anonymizes text by suppressing stylistic fingerprints
Researchers have developed a novel style-aware paraphrasing method to anonymize text, addressing the privacy risks posed by authorship attribution models. This approach utilizes large language models to create stylistic…
-
New annotation method improves Quran recitation ASR error detection
Researchers have developed a new method for annotating mistakes in Quran memorization transcripts generated by automatic speech recognition (ASR). This annotation process distinguishes between actual errors, repetitions…
-
New system enables real-time video understanding with VLMs
Researchers have developed a novel system designed for real-time video understanding using Vision-Language Models (VLMs). This system integrates lightweight clients with a server runtime that handles speech recognition,…
-
Xiaomi unveils LLM-based ASR for noisy environments
Researchers have introduced Xiaomi-CocktailASR-1, a novel end-to-end Automatic Speech Recognition (ASR) architecture designed to tackle the 'cocktail party problem' in multi-speaker environments. This LLM-based system u…
-
New evaluation method reveals ASR/audio LM struggles with English-Yoruba code-switching
A new research paper introduces a switch-aware evaluation method for automatic speech recognition (ASR) and audio language models (audio LMs) when processing code-switched speech, specifically focusing on English and Yo…
-
Eloquence team details multilingual QA approach for Interspeech 2026 challenge
The Eloquence team detailed their submission for the Interspeech 2026 MLC-SLM challenge, focusing on multilingual multiple-choice question answering across 21 languages. They explored three methods: fine-tuning Voxtral-…
-
AssemblyAI details voice agent load testing for production readiness
AssemblyAI has published a guide on load testing voice agents to ensure they perform reliably in production environments. The article emphasizes the critical difference between a controlled demo and real-world condition…
-
Olud Pulse tracks open-source AI adoption: Open Sora, Whisper, pgvector lead categories · 3 sources tracked
Olud Pulse, a tool that tracks the adoption of open-source AI, has released its latest scores. Open Sora leads in AI video generation, Whisper is highest in speech recognition and text-to-speech, and pgvector tops the v…
-
New framework audits bias and safety in voice AI customer care
A new research paper introduces a framework for auditing bias and safety in voice AI systems used for customer care. The proposed method categorizes voice AI architectures, including cascaded ASR-to-language model-to-TT…
-
New framework enhances speech anonymization while preserving data utility
Researchers have developed a new two-stage framework for speech anonymization that aims to preserve both linguistic content and acoustic identity while maintaining data utility. This framework uses a generative speech e…