发表机构
Huazhong University of Science and Technology; Zhejiang University(华中科技大学; 浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对3D动作训练数据不足的问题,提出MoVT框架,利用视频增强动作分词器丰富码本,并集成掩码Transformer,在多个关键指标上超越现有最先进方法。
AI 中文摘要
文本驱动的3D人体动作生成模型在响应多样且不受约束的文本提示时面临重大挑战,这主要是由于3D动作训练数据的可用性有限。为了解决这一问题,我们引入了MoVT,这是一个新颖的框架,能够有效利用广泛的人类动作视频来增强文本到动作的生成。我们方法的核心是跨模态增强动作分词器,它将离散的3D动作标记投影到2D域中。这种投影使我们能够利用从视频中提取的复杂、真实世界的动作模式来丰富动作码本。增强后的离散标记随后被映射回3D域,从而产生对齐的3D和2D码本,其表示复杂动作的能力得到增强。这些增强的码本被集成到一个生成式掩码Transformer中,该Transformer以模态无关的方式预测掩码动作标记索引。这使得能够利用从2D码本和带注释的动作视频生成的文本-索引对,进一步增强生成器。大量的实证评估表明,MoVT在多个关键指标上优于先前的最先进方法。
英文摘要
Text-driven 3D human motion generation models face significant challenges in responding to diverse and unconstrained textual prompts, primarily due to the limited availability of 3D motion training data. To address this, we introduce MoVT, a novel framework that effectively leverages the extensive range of human action videos to enhance text-to-motion generation. At the core of our approach is the cross-modal augmented motion tokenizer, which projects discrete 3D motion tokens into the 2D domain. This projection allows us to enrich the motion codebook with complex, real-world motion patterns derived from videos. The enriched discrete tokens are then mapped back to the 3D domain, resulting in aligned 3D and 2D codebooks with an enhanced capacity to represent intricate motions. These enhanced codebooks are integrated into a generative masked transformer, which predicts masked motion token indices in a modality-agnostic manner. This enables the use of text-index pairs, generated from the 2D codebook and annotated motion videos, to further enhance the generator. Extensive empirical evaluations show that MoVT performs favorably against prior state-of-the-art methods across multiple key metrics.