多模态受控连贯动作生成
Multi-Modal Controlled Coherent Motion Generation
浏览论文内容
中文总结 AI 辅助
本文提出MOCO,一种基于扩散的框架,通过解耦去噪步骤独立生成各模态动作并逐步协调,实现无需对齐数据的多模态连贯动作生成,在专用基准上优于现有方法。
中文摘要 AI 辅助
人类同时行走和说话是很自然的事情。本文解决了在由并发多模态输入(例如描述一个人行走并伴随语音音频的文本描述)驱动的3D虚拟形象动作生成中复制这种自然行为的挑战。现有方法受限于对齐多模态数据的稀缺性,通常通过顺序组合或加权求和的方式结合来自单个模态的动作。然而,它们往往导致不匹配或不真实的动作。为了克服这些限制,我们提出了MOCO,一种新颖的基于扩散的框架,能够处理多个同时输入,包括语音音频、文本描述和轨迹数据,以生成连贯且逼真的动作,而无需对齐的多模态数据。我们的关键创新在于解耦动作生成过程。在每个去噪步骤中,扩散模型从输入噪声中独立地为每个模态生成动作,并根据预定义的空间规则组装身体部位。然后,将得到的组合动作进行扩散,并作为后续去噪步骤的输入噪声。这种迭代方法使每个模态能够在整体动作的上下文中细化其贡献,逐步协调跨模态的动作。因此,生成的动作随着每次迭代变得越来越自然和流畅,实现连贯且同步的行为。我们使用专门构建的多模态基准来评估我们的方法。实验结果表明,MOCO优于现有基线,推进了3D虚拟形象多模态动作生成领域的发展。
英文摘要
It is natural for humans to walk and talk simultaneously. This paper tackles the challenge of replicating such natural behaviors in 3D avatar motion generation driven by concurrent multimodal inputs, such as a text description of a man walking alongside speech audio. Existing methods, constrained by the scarcity of aligned multimodal data, typically combine motions from individual modalities sequentially or through weighted sums. However, they often result in mismatched or unrealistic movements. To overcome these limitations, we propose MOCO, a novel diffusion-based framework capable of processing multiple simultaneous inputs, including speech audio, text descriptions, and trajectory data, to generate coherent and lifelike motions without requiring aligned multimodal data. Our key innovation lies in decoupling the motion generation process. During each denoising step, the diffusion model independently generates motions for each modality from the input noise and assembles the body parts according to predefined spatial rules. The resulting combined motion is then diffused and serves as the input noise for the subsequent denoising step. This iterative approach enables each modality to refine its contribution within the context of the overall motion, progressively harmonizing movements across modalities. Consequently, the generated motions become increasingly natural and fluid with each iteration, achieving coherent and synchronized behaviors. We evaluate our approach using a purpose-built multimodal benchmark. Experimental results demonstrate that MOCO outperforms existing baselines, advancing the field of multimodal motion generation for 3D avatars.
发表机构
- South China University of Technology(华南理工大学)
- Joy Future Academy(京东探索研究院)
- Peking University(北京大学)
机构由 AI 辅助整理,请以论文原文为准。