PulseAugur
实时 07:09:52
English(EN) UltraVoice: Scaling Fine-Grained Style-Controlled Speech Conversations for Spoken Dialogue Models

UltraVoice 数据集增强了对话模型中的语音风格控制

研究人员推出了 UltraVoice,这是一个旨在改进口语对话模型中细粒度语音风格控制的大规模数据集。该数据集包含超过 830 小时的语音对话,涵盖六个风格维度:情绪、语速、音量、口音、语言和复合风格的指令。在 UltraVoice 上微调 SLAM-OmniVocalNet 等模型,在风格可控性和指令遵循方面显示出显著的改进,同时不影响核心对话能力。该数据集的实用性也延伸到训练可控的文本到语音模型。 AI

影响 增强了口语对话系统中类似人类的交互,并改进了可控的文本到语音模型。

排序理由 该集群包含一篇详细介绍语音合成新数据集和方法的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

UltraVoice 数据集增强了对话模型中的语音风格控制

本文如何被排名

Signal score
24 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群包含一篇详细介绍语音合成新数据集和方法的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. arXiv cs.CL TIER_1 English(EN) · Wenming Tu, Guanrou Yang, Ruiqi Yan, Wenxi Chen, Ziyang Ma, Yipeng Kang, Kai Yu, Xie Chen, Zilong Zheng ·

    UltraVoice:为口语对话模型扩展细粒度风格控制语音对话

    arXiv:2510.22588v2 Announce Type: replace-cross Abstract: Spoken dialogue models currently lack the ability for fine-grained speech style control, a critical capability for human-like interaction that is often overlooked in favor of purely functional capabilities like reasoning a…