PulseAugur
EN
LIVE 16:32:21

New research tackles ASR challenges with synthetic speech, LLM optimization, and failure reduction

Researchers are developing advanced techniques to improve Automatic Speech Recognition (ASR) systems, particularly for challenging scenarios like code-switching and real-time applications. One paper proposes a code-mixing guided framework using synthetic speech to enhance ASR performance, reducing error rates on specific datasets. Another study introduces NIM4-ASR, an efficient and robust LLM-based ASR framework optimized for production, capable of handling noisy conditions and supporting large-scale customization. A third paper addresses catastrophic failures in neural-codec text-to-speech models, demonstrating that ASR self-verification and distillation can significantly reduce these errors, leading to more reliable speech synthesis. AI

IMPACT Advances in ASR and TTS aim to improve real-time applications, reduce errors in challenging speech scenarios, and enhance customization capabilities.

RANK_REASON Cluster consists of three academic papers on arXiv detailing advancements in speech recognition and synthesis technologies.

Read on arXiv cs.LG →

AI-generated summary · Google Gemini · from 3 sources. How we write summaries →

New research tackles ASR challenges with synthetic speech, LLM optimization, and failure reduction

COVERAGE [3]

  1. arXiv cs.AI TIER_1 English(EN) · Yue Heng Yeo, Haoyang Li, Yizhou Peng, Shreyas Gopal, Hexin Liu, Leibny Paola Garcia-Perera, Hardik B. Sailor, Jeremy H. M. Wong, Eng Siong Chng ·

    Improving Code-Switching ASR with Code-Mixing Guided Synthetic Speech

    arXiv:2606.19381v1 Announce Type: cross Abstract: Code-switch (CS) Automatic Speech Recognition (ASR) remains challenging due to limited availability of high quality CS text-speech pairs for training. Although synthetic data augmentation via Text-to-speech (TTS) has been explored…

  2. arXiv cs.CL TIER_1 English(EN) · Yuan Xie, Jiaqi Song, Guang Qiu, Xianliang Wang, Kai Qiao, Junfeng Yuan, Shengqing Liu, Yi Zhang, Bowen Chen, Ming Lei, Jie Gao, Jie Wu ·

    NIM4-ASR: Towards Efficient, Robust, and Customizable Real-Time LLM-Based ASR

    arXiv:2604.18105v2 Announce Type: replace-cross Abstract: Integrating large language models (LLMs) into automatic speech recognition (ASR) has become a mainstream paradigm in recent years. Although existing LLM-based ASR models demonstrate impressive performance on public benchma…

  3. arXiv cs.LG TIER_1 English(EN) · Ali Asaria, Tony Salomone, Deep Gandhi ·

    Reliable Neural-Codec Text-to-Speech by ASR Self-Verification and Distillation: Near-Zero Catastrophic Failures Across Models and Codecs

    arXiv:2606.18323v1 Announce Type: cross Abstract: Open autoregressive neural-codec text-to-speech (TTS) models sound excellent on typical inputs yet suffer stochastic catastrophic failures: on a meaningful fraction of utterances they emit silence, terminate early, or collapse int…