PulseAugur
实时 05:50:34
English(EN) Motion-Omni: End-to-End Joint Speech and Full-Body Motion for Spoken Dialogue

Motion-Omni 将语音和全身化身动作整合到一个框架中

研究人员开发了 Motion-Omni,一个将语音生成与全身化身动作相结合的端到端框架。与级联系统不同,Motion-Omni 允许语音和动作生成之间的联合优化,从而实现更协调、更自然的化身动作。该系统以 Qwen2.5-7B-Instruct 模型实例化,在保持低词错误率的同时,展示了更快的响应时间和在动作指标上的竞争力。 AI

影响 通过统一语音和动作生成,实现更自然、响应更快的化身交互。

排序理由 详细介绍用于联合语音和动作生成新框架的学术论文。 [lever_c_demoted from research: ic=1 ai=1.0]

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

Motion-Omni 将语音和全身化身动作整合到一个框架中

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
详细介绍用于联合语音和动作生成新框架的学术论文。 [lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
2 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
model release, paper, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
10 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+1 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准

报道来源 [2]

  1. Hugging Face Daily Papers TIER_1 English(EN) ·

    Motion-Omni:端到端联合语音与全身动作,用于口语对话

    Motion-Omni is an end-to-end framework that jointly generates spoken dialogue and full-body co-speech motion from shared hidden states, using scalable pseudo-labeling and a unified evaluation protocol to achieve real-time, aligned responses.

  2. arXiv cs.CV TIER_1 English(EN) · Chengqian Ma, Wei Tao, Haoyu Zhang, Yiwen Guo ·

    Motion-Omni: 语音对话的端到端联合语音与全身运动

    arXiv:2609.04250v1 Announce Type: cross Abstract: An avatar that holds a conversation should decide what to say and to move while saying it, yet these abilities live in separate model families: spoken dialogue models produce speech without motion, and co-speech motion models prod…