Automatic Speech Recognition
PulseAugur coverage of Automatic Speech Recognition — every cluster mentioning Automatic Speech Recognition across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
Voice AI platforms offer businesses a solution to customer service challenges
Voice AI platforms are emerging as a practical solution for businesses struggling with customer service wait times and high contact center turnover. These platforms utilize Automatic Speech Recognition (ASR), Natural La…
-
New research tackles LLM alignment, safety, and optimization challenges
Researchers are exploring new methods to improve the alignment and reliability of large language models (LLMs). One study identifies a vulnerability in byte-pair encoding (BPE) tokenization that can be exploited to bypa…
-
New research advances ASR for dysarthric speech and synthetic data use · 4 sources tracked
Researchers are exploring new methods to improve automatic speech recognition (ASR) systems. One study details how fine-tuning the Whisper model with personalized data significantly reduced word error rates for dysarthr…
-
Korean spoken QA research highlights ASR error impact on LLMs
A new research paper analyzes how errors in Korean speech recognition impact the performance of large language models (LLMs) in spoken question answering (SQA). The study found that the degradation caused by speech reco…
-
Speech representations impact 3D facial animation quality
Researchers have explored how different speech representations impact the quality of 3D facial animation. The study compared four families of speech representations, evaluating their effectiveness with two facial decode…
-
Open-source tools and ASR benchmarks advance local AI capabilities
This week's AI news highlights advancements in Automatic Speech Recognition (ASR) for bilingual voice agents and introduces two key open-source computer vision tools. The ASR focus is on benchmarking frontier models for…
-
ASR fine-tuned for Indian banking calls after 3-week effort
This article details the process of fine-tuning an Automatic Speech Recognition (ASR) system specifically for the unique challenges of Indian banking calls. The author spent three weeks experimenting with multiple model…
-
LLMs generate synthetic conversations to boost ASR training
Researchers have developed a novel method to enhance Automatic Speech Recognition (ASR) training for low-resource languages by generating synthetic conversational data. This pipeline uses LLMs to create dialogues, maps …
-
New ASR methods tackle compute scaling and multilingual evaluation
Researchers are developing new methods to improve automatic speech recognition (ASR) systems. One approach, LARM, uses a depth-conditioned looped Transformer to allow for adjustable test-time computation, achieving perf…
-
New Agentic ASR framework mimics human interaction for speech recognition
Researchers have introduced "Agentic ASR," a novel framework designed to improve automatic speech recognition (ASR) by mimicking human-like interactive correction. Unlike traditional single-pass systems, Agentic ASR ope…
-
New TARQ technique boosts ASR accuracy for rare words
Researchers have developed a new post-training quantization technique called TARQ, designed to improve the accuracy of Automatic Speech Recognition (ASR) systems, particularly for rare words. TARQ addresses a limitation…
-
Alibaba AI voice model ranks 5th globally, leads China on Speech Arena
Alibaba's new AI voice model, Fun-Realtime-TTS-Preview, has achieved a top global ranking on the Speech Arena benchmark, securing fifth place worldwide and first place in China. The model demonstrated strong performance…
-
Noisekit CLI generates realistic degraded audio for ASR benchmarking
A new command-line tool called noisekit has been released to help benchmark automatic speech recognition (ASR) systems. It generates realistic degraded audio datasets by applying various noise and distortion conditions …
-
Intel NPU accelerates smart home ASR, outperforming CPU on speed and energy
A user has successfully utilized their Intel Arrow Lake NPU for Automatic Speech Recognition (ASR) in a smart home setup, achieving significant performance gains. The NPU processed a 10-second audio clip 4.8 times faste…
-
New method enhances spoken dialogue systems by diagnosing ASR-LLM errors
Researchers have developed a novel approach to improve spoken dialogue systems by addressing error propagation in cascaded Automatic Speech Recognition (ASR) and Large Language Model (LLM) pipelines. This new method use…
-
AI voice assistants in 2026 offer advanced capabilities for personal and business use
AI voice assistants in 2026 are significantly more advanced, leveraging LLMs, ASR, ML, and NLP to understand natural speech, learn continuously, and personalize responses. These assistants are categorized into personal …
-
New neural layer nASR enhances EEG artifact removal for BCIs
Researchers have developed nASR, a novel trainable neural layer designed to improve Electroencephalogram (EEG) signal processing for Brain-Computer Interfaces (BCIs). This new layer addresses limitations in existing Art…
-
Voice AI paradox: Advanced chat, basic failures
Voice AI assistants like Yandex's Alisa exhibit a paradox of advanced conversational abilities alongside basic functional failures, stemming from their complex architecture. This hybrid system combines speech recognitio…
-
Sakana AI's KAME architecture injects LLM knowledge into speech AI without latency
Sakana AI has developed KAME, a novel tandem architecture for speech-to-speech AI that aims to combine the speed of direct systems with the knowledge depth of LLM-based approaches. KAME operates with two asynchronous co…
-
Tamazight single-speaker speech dataset released on Hugging Face
A new single-speaker speech dataset for the Tamazight language has been released on Hugging Face and the Mozilla Data Collective. This dataset is intended for use in AI applications such as automatic speech recognition …