PulseAugur
实时 10:11:20
English(EN) StreamTalk: Streaming Co-Speech Gesture Generation with Key-Pose Anchoring

StreamTalk 框架以更高的准确性生成实时语音同步手势

研究人员开发了 StreamTalk,一个用于实时生成逼真语音同步手势的新框架。与之前存在长序列累积漂移问题的开环方法不同,StreamTalk 采用闭环系统,包含生成-检索-精炼循环。该方法使用关键姿态作为锚点,以限制漂移并提高轨迹精度。该框架还采用了随机姿态掩码和感知部件的 DiT 等技术,以增强运动恢复并减少不同运动流之间的干扰,在 BEAT2 数据集上取得了最先进的成果。 AI

影响 这项研究可能带来更自然、更具吸引力的虚拟化身和人机交互。

排序理由 该集群包含一篇详细介绍手势生成新方法的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CV 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

StreamTalk 框架以更高的准确性生成实时语音同步手势

报道来源 [1]

  1. arXiv cs.CV TIER_1 English(EN) · Xiangyue Zhang, Jianfang Li, Jiaxu Zhang, Kaixing Yang, Steven Hoi ·

    StreamTalk: Streaming Co-Speech Gesture Generation with Key-Pose Anchoring

    arXiv:2608.01643v1 Announce Type: new Abstract: Real-time co-speech gesture generation must produce 3D motion clip by clip as speech arrives. Existing streaming methods are open-loop: each clip depends on past context, but the model cannot check or correct its trajectory. Small e…