PulseAugur
EN
LIVE 08:17:09

New discrete diffusion model enhances speech inpainting and editing

Researchers have developed SIEDD, a novel discrete diffusion model designed for speech inpainting and editing. This framework, named HiCoDD, operates on hierarchical codec tokens, iteratively refining masked segments while maintaining speaker identity, prosody, and recording conditions. SIEDD demonstrates superior performance on the RealEdit benchmark for speech editing and outperforms autoregressive baselines in speech inpainting tasks, showcasing the benefits of explicitly modeling codec hierarchies for context-preserving audio reconstruction. AI

IMPACT This model could improve audio editing tools and content creation by enabling more precise reconstruction and modification of speech segments.

RANK_REASON The cluster contains a research paper detailing a new model and its performance on benchmarks. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.CL →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New discrete diffusion model enhances speech inpainting and editing

COVERAGE [1]

  1. arXiv cs.CL TIER_1 English(EN) · Iftach Shoham, Tali Dror, Oren Gal, Haim Permuter, Gilad Katz, Eliya Nachmani ·

    Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing

    arXiv:2608.06424v1 Announce Type: cross Abstract: Speech recordings often contain missing, corrupted, or incorrect regions that must be reconstructed or modified without re-synthesizing the entire utterance. Speech inpainting restores missing segments, whereas speech editing repl…