PulseAugur
EN
LIVE 08:10:32

New ASR method enables controllable verbatim transcription and timing

Researchers have developed a new method to control the transcription style of Automatic Speech Recognition (ASR) models, addressing issues of decoding instability and evaluation confounding caused by inconsistent transcription styles. By training models with coverage-aware decoder task tokens on parallel verbatim and intended transcript pairs, they achieved significant improvements in disfluency detection, even in languages not used during English-only training. The approach also enhances word-level timing accuracy and introduces a new task, 'verbatimize', for scalable creation of high-quality verbatim transcriptions. AI

IMPACT This research could lead to more accurate and reliable speech-to-text systems, improving user experience and data analysis in various applications.

RANK_REASON The item is an academic paper detailing a new method for ASR models. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.CL →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New ASR method enables controllable verbatim transcription and timing

COVERAGE [1]

  1. arXiv cs.CL TIER_1 English(EN) · Laurin Wagner (nyra labs), Mario Zusag (nyra labs), Bernhard Thallinger (nyra labs) ·

    Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level Timing

    arXiv:2607.18934v1 Announce Type: new Abstract: Modern ASR models trained on heterogeneously annotated data treat transcription style (verbatim vs. intended) as an uncontrolled latent variable, causing measurable decoding instability, evaluation confounding (up to 60% of reported…