PulseAugur
实时 08:17:37
English(EN) EMODY Flow: Emotion-Aware Audio-Driven Full-Body Motion Generation

EMODY Flow 从音频生成情感感知的全身运动

研究人员开发了EMODY Flow,一个用于生成与语音和情感线索同步的全身运动的新颖框架。该系统解决了现有模型中情感条件化经常被低估的局限性。EMODY Flow是一个拥有约3500万参数的轻量级模型,与冻结的Qwen-3 Omni模型集成,并利用其Mimi音频编解码器。它采用两个并行的Diffusion Transformer (DiT) 生成器来生成身体姿势和面部表情,并通过一个辅助情感分类器进行增强,以确保生成运动中独特的情感表达。该框架在BEAT2数据集上实现了手势质量、相关性和多样性的最先进结果,显著优于以前的方法。 AI

影响 通过实现由音频驱动的更具表现力和情感共鸣的角色动画,推动了具身AI的发展。

排序理由 学术论文,详细介绍了一个新模型及其在基准测试上的表现。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.LG 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

EMODY Flow 从音频生成情感感知的全身运动

本文如何被排名

Signal score
18 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
学术论文,详细介绍了一个新模型及其在基准测试上的表现。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. arXiv cs.LG TIER_1 English(EN) · Harsh Kumar Agarwal, Xavier Alameda-Pineda, Olivier Perrotin ·

    EMODY Flow:情感感知音频驱动的全身运动生成

    arXiv:2609.16011v1 Announce Type: cross Abstract: Embodied conversational agents require synchronized full-body motion (body gestures and facial expressions) that aligns with speech and emotional state. Omni-modal large language models excel at multimodal understanding but produc…