speech recognition
PulseAugur coverage of speech recognition — every cluster mentioning speech recognition across labs, papers, and developer communities, ranked by signal.
- instance of alphaXiv 90%
- used by alphaXiv 70%
- used by word error rate 70%
- used by LibriSpeech 70%
- used by whisper tiny 70%
- used by Vad 70%
- competes with speech synthesis 60%
- used by Gotit.pub 60%
- instance of word error rate 60%
- instance of ScienceCast 60%
- affiliated with speech synthesis 50%
- other word error rate 50%
17 day(s) with sentiment data
-
AssemblyAI: Voice agent success hinges on STT accuracy, not flash
AssemblyAI argues that the most crucial factor for a successful voice agent is the accuracy of its speech-to-text (STT) foundation, rather than superficial metrics like speed or dashboard aesthetics. The company emphasi…
-
New DoS attack targets end-to-end speech LLMs with acoustic perturbations
Researchers have developed a new denial-of-service (DoS) attack specifically targeting end-to-end (E2E) speech large language models (LLMs). Unlike previous attacks that relied on text prompt manipulation, this method i…
-
VoxZip framework slashes audio LLM KV cache needs by 20x
Researchers have developed VoxZip, a novel two-stage framework designed to compress the KV cache for long-context audio inference in Speech Large Language Models. This method uses Automatic Speech Recognition (ASR) tran…
-
Local AI Updates: llama.cpp, PyTorch, Kimi-K3, and NVIDIA NeMo Speech 3.0
Recent updates in the local AI and open-source model space include performance enhancements for llama.cpp with CUDA fusion, addressing critical quantization bugs in PyTorch for AMD GPUs, and the trending Moonshot AI Kim…
-
Developer creates local voice input extension for Pi using nemotron 3.5 ASR
A developer has created a lightweight, local voice input extension for the Pi coding terminal, utilizing NVIDIA's nemotron 3.5 0.6B Automatic Speech Recognition model. This extension aims for simplicity, running an STT …
-
New framework aims to decolonize automated speech recognition systems
This paper introduces a framework for developing more culturally competent automated speech recognition (ASR) systems. It argues that current ASR failures with low-resource and Indigenous languages are not just technica…
-
New research compares context biasing and speech LLMs for rare word recognition in ASR
A new research paper published on arXiv explores methods for improving automatic speech recognition (ASR) systems' ability to recognize new and rare words. The study compares context biasing techniques, which supply a w…
-
Challenges in collecting high-quality speech and video datasets for AI
Collecting high-quality datasets for multimodal AI, specifically studio-quality speech and egocentric video, presents significant challenges. These include maintaining consistent recording environments, managing device …
-
New distillation method enhances multilingual ASR systems
Researchers have developed a new method called Language-Specialized Multi-Teacher On-Policy Distillation (LS-MOPD) to improve multilingual Automatic Speech Recognition (ASR) systems. This approach addresses optimization…
-
Study probes ASR model for dysarthric speech effects
Researchers have conducted a layer-wise probing analysis on a transformer Automatic Speech Recognition (ASR) model using Mandarin dysarthric speech. The study found that phoneme boundary information remains weak across …
-
Speech2Grasp framework enables humanoid robots to grasp objects via spoken commands
Researchers have developed Speech2Grasp, a novel framework that enables humanoid robots to understand and act upon spoken commands for grasping objects. This approach efficiently transfers capabilities from existing tex…
-
New PARSE framework enhances concept erasure in diffusion models
Researchers have developed a new training-free framework called PARSE (Preservation-aware Adaptive Ranked Subspace Expansion) designed to improve concept erasure in text-to-image diffusion models. Existing methods often…
-
AI framework converts emergency voice calls into structured data
Researchers have developed a new AI framework called SIREN that can process voice communications from emergency responders to create structured, machine-readable information. This framework integrates automatic speech r…
-
New benchmark dataset targets Indian languages for speech technology
Researchers have developed Indic DiarBench, a new benchmark dataset designed to improve speech technology for the 22 scheduled languages of India. This dataset includes approximately 108 hours of audio from various real…
-
audio37 launches TTS and voice cloning, seeks feature ideas
The audio37 tool has been released, offering Text-to-Speech (TTS) capabilities along with voice cloning. The developer is soliciting community input on potential future features, such as speech recognition (STT) transcr…
-
Voice cloning enhances paralinguistic tasks and cross-lingual clinical speech analysis
A new research paper explores the use of voice cloning for data augmentation in paralinguistic tasks, particularly for clinical applications where labeled data is scarce. The study benchmarks eight voice cloning models,…
-
New MEUSLI projector enables multilingual ASR and speech understanding
Researchers have developed MEUSLI, a novel multilingual projector designed to link speech encoders with large language models (LLMs) for advanced speech processing tasks. This system extends existing monolingual project…
-
Voice cloning models preserve paralinguistic signals for clinical speech tasks
Researchers have evaluated eight voice cloning models to determine their effectiveness in preserving paralinguistic signals for speech tasks, particularly in clinical settings where labeled data is scarce. The study fou…
-
AssemblyAI details speech-to-text API fundamentals
AssemblyAI has released a guide detailing the fundamentals of its speech-to-text API. The guide explains how to authenticate requests using an API key, submit audio files for transcription asynchronously, and poll for j…
-
AssemblyAI explains advanced AI voice recognition capabilities
AssemblyAI has detailed how modern AI voice recognition technology goes beyond simple speech-to-text conversion. The technology now incorporates machine learning models capable of identifying speakers, detecting sentime…