Researchers have developed GEPARD, a novel text-to-speech model designed for real-time dialogue applications. This model utilizes an LLM backbone for autoregressive speech generation and a neural codec for waveform decoding, enabling it to stream audio incrementally as text is processed. GEPARD is engineered to operate efficiently with standard LLM serving engines like vLLM, achieving a real-time factor of approximately 0.067 (15x faster than real-time) for single streams and significant aggregate speedups under concurrent usage. AI
IMPACT This model could significantly improve the responsiveness and naturalness of AI-powered conversational agents.
RANK_REASON Research paper detailing a new TTS model. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →