A new speech restoration model named Diamond has been released by nineninesix.ai. This autoregressive sequence-to-sequence model utilizes a two-transformer architecture, including a time-transformer and a compact depth-transformer, to reconstruct degraded audio into high-fidelity 44.1 kHz speech. Diamond achieves competitive performance against other open-source restoration models, even when trained from scratch with a relatively small parameter count and training time. AI
IMPACT Offers improved speech restoration capabilities, potentially impacting audio production and dataset cleansing.
RANK_REASON Release of a new model with technical details and benchmark results. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Hugging Face Trending Models →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →