用于人体动作生成的时空解耦自回归扩散模型
Spatiotemporally Decoupled Autoregressive Diffusion Model for Human Motion Generation
- AI Institute, Shanghai Jiao Tong University(上海交通大学AI学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出名为DeMoDiff的时空解耦框架,通过针对每个身体关节编码的时空VAE与融入时空掩码和注意力的自回归扩散生成器,在HumanML3D和KIT-ML数据集上取得最优重构性能与动作生成效果,且具备强时空编辑能力。
AI中文摘要:
文本驱动的人体动作合成已取得显著发展,核心包含动作表示与生成架构两大模块。在表示层面,基于向量量化(VQ)的方法将动作数据压缩为离散 token,而基于隐空间的模型则直接在连续空间中操作,但两类表示均存在明显局限:基于VQ的方法存在固有的信息损失,会损害生成动作的质量、多样性与泛化能力;针对整体全身动作的连续表示则会阻碍身体部位级别的灵活性。在架构层面,扩散模型与自回归扩散模型已展现出优越性,但对单个身体部位的细粒度可控性仍有限。因此,本文提出名为DeMoDiff的统一时空解耦框架,对表示与架构进行联合重新设计:为提升表示提取能力并提供更强的部位级可控性,本文提出一种时空VAE,对每个身体关节进行编码,而非将全身动作压缩至单一隐空间;随后,本文将时空掩码与注意力机制融入自回归扩散生成器,兼顾生成能力与可控编辑性。在HumanML3D与KIT-ML数据集上开展的大量实验表明,本文模型达到了最先进的重构性能与引人注目的动作生成结果,且框架展现出强大的时空编辑能力,进一步验证了其有效性。项目页面为:this https URL
英文摘要:
Text-driven human motion synthesis has made substantial development with two core modules of motion representation and generative architecture. For representation, Vector Quantization (VQ)-based methods compress motion data into discrete tokens while latent-based models operate directly in continuous space. However, both of these representations exhibit significant limitations. VQ-based methods suffer from inherent information loss, which compromises the quality, diversity, and generalization of generated motions, while continuous representation on holistic whole-body motion hinders part-level flexibility. For architecture, diffusion and autoregressive diffusion models have demonstrated their superiority, yet the fine-grained controllability over individual body parts is also limited. Thus, we propose a unified spatiotemporally decoupled framework named DeMoDiff, which jointly redesigns representation and architecture. To enhance representation extraction capabilities and offer greater part-level controllability, we present a spatial-temporal VAE that encodes each body joint rather than compressing the whole-body motion into a single latent space. Then, we incorporate spatial-temporal masking and attention mechanisms into an autoregressive diffusion generator, achieving both generative capability and controllable editability. Extensive experiments on the HumanML3D and KIT-ML datasets demonstrate that our model achieves state-of-the-art reconstruction performance and compelling motion generation results. Moreover, our framework demonstrates strong temporal and spatial editing capabilities, further validating its effectiveness. Our project page: https://rex0191.github.io/DeMoDiff/