Researchers have developed a new method to control the transcription style of Automatic Speech Recognition (ASR) models, addressing issues of decoding instability and evaluation confounding caused by inconsistent transcription styles. By training models with coverage-aware decoder task tokens on parallel verbatim and intended transcript pairs, they achieved significant improvements in disfluency detection, even in languages not used during English-only training. The approach also enhances word-level timing accuracy and introduces a new task, 'verbatimize', for scalable creation of high-quality verbatim transcriptions. AI
IMPACT This research could lead to more accurate and reliable speech-to-text systems, improving user experience and data analysis in various applications.
RANK_REASON The item is an academic paper detailing a new method for ASR models. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →