PulseAugur
EN
LIVE 10:59:43

New X2Streaming-TTS enables true token-level speech synthesis from streaming text

Researchers have developed X2Streaming-TTS, a novel text-to-speech system designed for true token-level synthesis from streaming text. This framework processes text tokens as they arrive, generating speech without needing future input. It incorporates causal commitment to manage uncertain text prefixes and causal speech-state inheritance to maintain acoustic continuity across segment boundaries. Experiments indicate that X2Streaming-TTS surpasses existing pseudo-streaming models in quality and achieves low latency, with a median time to first audio token of 15.8 ms for single requests. AI

IMPACT Enables lower-latency spoken dialogue systems by providing true token-level speech synthesis.

RANK_REASON Academic paper detailing a new model and its capabilities. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.CL →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New X2Streaming-TTS enables true token-level speech synthesis from streaming text

COVERAGE [1]

  1. arXiv cs.CL TIER_1 English(EN) · Rime Wen, Zehan Liu, Shawn Qin, Lights Shi, Roy Gan, Hao Wang, Qian Wang ·

    X2Streaming-TTS: Causal Token-Level Text-to-Speech from Streaming Text with Speech-State Inheritance

    arXiv:2608.18661v1 Announce Type: new Abstract: Streaming text-to-speech is essential for low-latency spoken dialogue systems, yet many systems wait for sentence-level text and are therefore only pseudo-streaming. True token-level synthesis must generate speech from uncertain pre…