PulseAugur
实时 06:51:40
English(EN) GEPARD - Generative, Prosody-aware, Autoregressive text-to-speech model for Realtime Dialogue

新的 GEPARD TTS 模型为对话实现 15 倍实时速度

研究人员开发了 GEPARD,这是一种新颖的文本到语音模型,专为实时对话应用而设计。该模型利用 LLM 主干进行自回归语音生成,并利用神经编解码器进行波形解码,使其能够在处理文本时逐步流式传输音频。GEPARD 经过工程设计,可与 vLLM 等标准 LLM 服务引擎高效运行,在单流情况下实现约 0.067 的实时因子(比实时快 15 倍),并在并发使用下实现显著的聚合加速。 AI

影响 该模型可以显著提高 AI 驱动的对话代理的响应能力和自然度。

排序理由 详细介绍新型 TTS 模型的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的 GEPARD TTS 模型为对话实现 15 倍实时速度

本文如何被排名

Signal score
27 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
详细介绍新型 TTS 模型的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
model release, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. arXiv cs.CL TIER_1 English(EN) · Denis Pavlov, Ulanbek Abdurazakov, Nursultan Bakashov ·

    GEPARD - 面向实时对话的生成式、韵律感知、自回归文本到语音模型

    arXiv:2609.04222v1 Announce Type: cross Abstract: We present GEPARD (Generative, Prosody-aware, Autoregressive text-to-speech model for Realtime Dialogue), a streaming text-to-speech model for real-time spoken dialogue. GEPARD generates speech autoregressively with an LLM backbon…