arXiv:2609.38400v1 Announce Type: cross Abstract: Co-speech gestures for robots must adapt not only to speech and embodiment, but also to the workspace available for performing the motion. Since the same speech can be accompanied by different gestures, a robot can respond to work…
arXiv:2609.39575v1 Announce Type: cross Abstract: Generating full-body co-speech motion for humanoid robots requires coordinating speech prosody, linguistic content, and embodiment-specific motion. To this end, we present ECHO-G, a framework jointly conditioned on speech audio an…
Generating full-body co-speech motion for humanoid robots requires coordinating speech prosody, linguistic content, and embodiment-specific motion. To this end, we present ECHO-G, a framework jointly conditioned on speech audio and timed transcripts. Its Speech-Grounded Diffusion…
arXiv cs.AI
TIER_1English(EN)·Yuanzhuo Hu, Zehan Liu, Xiaoyi Qin, Ming Li·
arXiv:2609.36624v1 Announce Type: new Abstract: Joint speech-gesture synthesis must coordinate two modalities despite limited paired data. Existing approaches often lack bidirectional interaction, have limited language coverage, or simplify body and finger representations. We pre…
arXiv cs.CV
TIER_1English(EN)·Zhirui Xing, Long Ye, Kaige Li, Ziyi Xu, Ming Meng·
arXiv:2609.36685v1 Announce Type: new Abstract: Co-speech gesture generation aims to synthesize natural gestures that are both temporally synchronized with speech and semantically consistent with the spoken content. Although recent methods can generate rhythmically plausible moti…