PulseAugur
EN
LIVE 10:41:02

New AI models enhance co-speech gesture generation with improved realism and efficiency · 4 sources tracked

Researchers have developed several new methods for generating realistic co-speech gestures. SemTalk focuses on integrating base motions with semantic gestures by learning them separately and adaptively fusing them. GestureLSM addresses challenges in quality and speed by modeling spatial-temporal interactions between body regions and using flow matching for efficient sampling. GlobalDiff mitigates error accumulation in long-horizon gesture generation by operating directly on global joint rotations and incorporating multi-level constraints. EchoMask utilizes a speech-queried attention mechanism within a masked modeling framework to guide gesture generation by selectively masking frames based on speech cues. AI

IMPACT These advancements could lead to more realistic and responsive digital avatars and embodied agents in virtual environments and real-time applications.

RANK_REASON Multiple research papers detailing new methods for co-speech motion generation.

Read on arXiv cs.CV →

AI-generated summary · Google Gemini · from 4 sources. How we write summaries →

New AI models enhance co-speech gesture generation with improved realism and efficiency · 4 sources tracked

COVERAGE [4]

  1. arXiv cs.CV TIER_1 English(EN) · Xiangyue Zhang, Jianfang Li, Jiaxu Zhang, Ziqiang Dang, Jianqiang Ren, Liefeng Bo, Zhigang Tu ·

    SemTalk: Holistic Co-speech Motion Generation with Frame-level Semantic Emphasis

    arXiv:2412.16563v4 Announce Type: replace Abstract: Co-speech gesture generation must carefully integrate common rhythmic motion with rare yet essential semantic gestures. In this work, we propose SemTalk for holistic co-speech gesture generation with frame-level semantic emphasi…

  2. arXiv cs.CV TIER_1 English(EN) · Pinxin Liu, Luchuan Song, Junhua Huang, Haiyang Liu, Junfan Zhu, Chenliang Xu ·

    GestureLSM: Latent Shortcut based Co-Speech Gesture Generation with Spatial-Temporal Modeling

    arXiv:2501.18898v4 Announce Type: replace Abstract: Generating full-body human gestures based on speech signals remains challenges on quality and speed. Existing approaches model different body regions such as body, legs and hands separately, which fail to capture the spatial int…

  3. arXiv cs.CV TIER_1 English(EN) · Xiangyue Zhang, Jianfang Li, Jianqiang Ren, Jiaxu Zhang ·

    Mitigating Error Accumulation in Co-Speech Motion Generation via Global Rotation Diffusion and Multi-Level Constraints

    arXiv:2511.10076v3 Announce Type: replace Abstract: Reliable long-horizon co-speech gesture generation requires precise motion representation and consistent structural priors across all joints. Existing generative methods typically operate on local joint rotations, which are defi…

  4. arXiv cs.CV TIER_1 English(EN) · Xiangyue Zhang, Jianfang Li, Jiaxu Zhang, Jianqiang Ren, Liefeng Bo, Zhigang Tu ·

    EchoMask: Speech-Queried Attention-based Mask Modeling for Holistic Co-Speech Motion Generation

    arXiv:2504.09209v3 Announce Type: replace-cross Abstract: Masked modeling has shown promise in co-speech gesture generation. However, it struggles to identify semantically significant frames for effective motion masking. In this work, we propose a speech-queried attention-based m…