PulseAugur
实时 11:40:49
English(EN) SemTalk: Holistic Co-speech Motion Generation with Frame-level Semantic Emphasis

新的AI模型通过提高真实感和效率来增强同声手势生成 · 跟踪4个来源

研究人员开发了几种生成逼真同声手势的新方法。SemTalk通过分别学习基础动作和语义手势,然后自适应地融合它们,来专注于整合这两者。GestureLSM通过对身体区域之间的时空交互进行建模并使用流匹配进行高效采样,来解决质量和速度方面的挑战。GlobalDiff通过直接对全局关节旋转进行操作并纳入多级约束,来减轻长时序手势生成中的误差累积。EchoMask在掩码建模框架内利用语音查询的注意力机制,通过根据语音线索选择性地掩码帧来指导手势生成。 AI

影响 这些进步可能为虚拟环境和实时应用中的数字虚拟形象和具身智能体带来更逼真、更具响应性的体验。

排序理由 多篇研究论文详细介绍了同声运动生成的新方法。

在 arXiv cs.CV 阅读 →

AI 生成摘要 · Google Gemini · 来自 4 个来源。 我们如何撰写摘要 →

新的AI模型通过提高真实感和效率来增强同声手势生成 · 跟踪4个来源

报道来源 [4]

  1. arXiv cs.CV TIER_1 English(EN) · Xiangyue Zhang, Jianfang Li, Jiaxu Zhang, Ziqiang Dang, Jianqiang Ren, Liefeng Bo, Zhigang Tu ·

    SemTalk:具有帧级语义强调的整体同说话动作生成

    arXiv:2412.16563v4 Announce Type: replace Abstract: Co-speech gesture generation must carefully integrate common rhythmic motion with rare yet essential semantic gestures. In this work, we propose SemTalk for holistic co-speech gesture generation with frame-level semantic emphasi…

  2. arXiv cs.CV TIER_1 English(EN) · Pinxin Liu, Luchuan Song, Junhua Huang, Haiyang Liu, Junfan Zhu, Chenliang Xu ·

    GestureLSM:基于潜在捷径的共语手势生成与时空建模

    arXiv:2501.18898v4 Announce Type: replace Abstract: Generating full-body human gestures based on speech signals remains challenges on quality and speed. Existing approaches model different body regions such as body, legs and hands separately, which fail to capture the spatial int…

  3. arXiv cs.CV TIER_1 English(EN) · Xiangyue Zhang, Jianfang Li, Jianqiang Ren, Jiaxu Zhang ·

    通过全局旋转扩散和多级约束缓解共语运动生成中的误差累积

    arXiv:2511.10076v3 Announce Type: replace Abstract: Reliable long-horizon co-speech gesture generation requires precise motion representation and consistent structural priors across all joints. Existing generative methods typically operate on local joint rotations, which are defi…

  4. arXiv cs.CV TIER_1 English(EN) · Xiangyue Zhang, Jianfang Li, Jiaxu Zhang, Jianqiang Ren, Liefeng Bo, Zhigang Tu ·

    EchoMask:基于语音查询的注意力掩码建模,用于整体式同声运动生成

    arXiv:2504.09209v3 Announce Type: replace-cross Abstract: Masked modeling has shown promise in co-speech gesture generation. However, it struggles to identify semantically significant frames for effective motion masking. In this work, we propose a speech-queried attention-based m…