PulseAugur
EN
LIVE 19:16:04
ENTITY speech recognition

speech recognition

PulseAugur coverage of speech recognition — every cluster mentioning speech recognition across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
29
104 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
17
80 over 90d
TIER MIX · 90D
TOPICS
RELATIONSHIPS
SENTIMENT · 30D

17 day(s) with sentiment data

RECENT · PAGE 1/6 · 104 TOTAL
  1. COMMENTARY · CL_197186 ·

    AssemblyAI: Voice agent success hinges on STT accuracy, not flash

    AssemblyAI argues that the most crucial factor for a successful voice agent is the accuracy of its speech-to-text (STT) foundation, rather than superficial metrics like speed or dashboard aesthetics. The company emphasi…

  2. TOOL · CL_195982 ·

    New DoS attack targets end-to-end speech LLMs with acoustic perturbations

    Researchers have developed a new denial-of-service (DoS) attack specifically targeting end-to-end (E2E) speech large language models (LLMs). Unlike previous attacks that relied on text prompt manipulation, this method i…

  3. TOOL · CL_193340 ·

    VoxZip framework slashes audio LLM KV cache needs by 20x

    Researchers have developed VoxZip, a novel two-stage framework designed to compress the KV cache for long-context audio inference in Speech Large Language Models. This method uses Automatic Speech Recognition (ASR) tran…

  4. TOOL · CL_189120 ·

    Local AI Updates: llama.cpp, PyTorch, Kimi-K3, and NVIDIA NeMo Speech 3.0

    Recent updates in the local AI and open-source model space include performance enhancements for llama.cpp with CUDA fusion, addressing critical quantization bugs in PyTorch for AMD GPUs, and the trending Moonshot AI Kim…

  5. TOOL · CL_187608 ·

    Developer creates local voice input extension for Pi using nemotron 3.5 ASR

    A developer has created a lightweight, local voice input extension for the Pi coding terminal, utilizing NVIDIA's nemotron 3.5 0.6B Automatic Speech Recognition model. This extension aims for simplicity, running an STT …

  6. TOOL · CL_187373 ·

    New framework aims to decolonize automated speech recognition systems

    This paper introduces a framework for developing more culturally competent automated speech recognition (ASR) systems. It argues that current ASR failures with low-resource and Indigenous languages are not just technica…

  7. TOOL · CL_187366 ·

    New research compares context biasing and speech LLMs for rare word recognition in ASR

    A new research paper published on arXiv explores methods for improving automatic speech recognition (ASR) systems' ability to recognize new and rare words. The study compares context biasing techniques, which supply a w…

  8. COMMENTARY · CL_187825 ·

    Challenges in collecting high-quality speech and video datasets for AI

    Collecting high-quality datasets for multimodal AI, specifically studio-quality speech and egocentric video, presents significant challenges. These include maintaining consistent recording environments, managing device …

  9. TOOL · CL_183272 ·

    New distillation method enhances multilingual ASR systems

    Researchers have developed a new method called Language-Specialized Multi-Teacher On-Policy Distillation (LS-MOPD) to improve multilingual Automatic Speech Recognition (ASR) systems. This approach addresses optimization…

  10. TOOL · CL_180522 ·

    Study probes ASR model for dysarthric speech effects

    Researchers have conducted a layer-wise probing analysis on a transformer Automatic Speech Recognition (ASR) model using Mandarin dysarthric speech. The study found that phoneme boundary information remains weak across …

  11. TOOL · CL_172047 ·

    Speech2Grasp framework enables humanoid robots to grasp objects via spoken commands

    Researchers have developed Speech2Grasp, a novel framework that enables humanoid robots to understand and act upon spoken commands for grasping objects. This approach efficiently transfers capabilities from existing tex…

  12. TOOL · CL_167703 ·

    New PARSE framework enhances concept erasure in diffusion models

    Researchers have developed a new training-free framework called PARSE (Preservation-aware Adaptive Ranked Subspace Expansion) designed to improve concept erasure in text-to-image diffusion models. Existing methods often…

  13. TOOL · CL_167501 ·

    AI framework converts emergency voice calls into structured data

    Researchers have developed a new AI framework called SIREN that can process voice communications from emergency responders to create structured, machine-readable information. This framework integrates automatic speech r…

  14. TOOL · CL_167397 ·

    New benchmark dataset targets Indian languages for speech technology

    Researchers have developed Indic DiarBench, a new benchmark dataset designed to improve speech technology for the 22 scheduled languages of India. This dataset includes approximately 108 hours of audio from various real…

  15. TOOL · CL_165504 ·

    audio37 launches TTS and voice cloning, seeks feature ideas

    The audio37 tool has been released, offering Text-to-Speech (TTS) capabilities along with voice cloning. The developer is soliciting community input on potential future features, such as speech recognition (STT) transcr…

  16. TOOL · CL_165096 ·

    Voice cloning enhances paralinguistic tasks and cross-lingual clinical speech analysis

    A new research paper explores the use of voice cloning for data augmentation in paralinguistic tasks, particularly for clinical applications where labeled data is scarce. The study benchmarks eight voice cloning models,…

  17. RESEARCH · CL_165028 ·

    New MEUSLI projector enables multilingual ASR and speech understanding

    Researchers have developed MEUSLI, a novel multilingual projector designed to link speech encoders with large language models (LLMs) for advanced speech processing tasks. This system extends existing monolingual project…

  18. TOOL · CL_175926 ·

    Voice cloning models preserve paralinguistic signals for clinical speech tasks

    Researchers have evaluated eight voice cloning models to determine their effectiveness in preserving paralinguistic signals for speech tasks, particularly in clinical settings where labeled data is scarce. The study fou…

  19. TOOL · CL_157935 ·

    AssemblyAI details speech-to-text API fundamentals

    AssemblyAI has released a guide detailing the fundamentals of its speech-to-text API. The guide explains how to authenticate requests using an API key, submit audio files for transcription asynchronously, and poll for j…

  20. TOOL · CL_155782 ·

    AssemblyAI explains advanced AI voice recognition capabilities

    AssemblyAI has detailed how modern AI voice recognition technology goes beyond simple speech-to-text conversion. The technology now incorporates machine learning models capable of identifying speakers, detecting sentime…