发表机构
Southeast University; National University of Singapore; Opus AI(东南大学; 新加坡国立大学; Opus AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对文本-运动双向生成中自回归模型顺序固定与误差累积的问题,提出统一掩码离散扩散框架BiMoGen,采用解耦训练与生成感知自校正,在HumanML3D和KIT-ML上取得竞争性能。
AI 中文摘要
文本到运动生成和运动到文本描述是人体运动建模中的两个基本任务,两者都基于相同的底层运动-文本对应关系。现有的统一方法大多依赖于自回归建模,这施加了固定的生成顺序,因此不适合语言和运动之间的双向依赖,使得早期预测错误作为固定上下文持续存在,并降低时间连贯性和跨模态一致性。掩码离散扩散通过迭代双向预测对序列进行建模,提供了一种自然的补救措施。因此,我们提出了BiMoGen(双向运动-文本生成),一个用于双向运动-文本建模的统一掩码离散扩散框架。为了稳定训练,我们设计了解耦的单模态和跨模态训练,其中掩码预训练首先在配对的运动-文本序列上建立跨模态对应关系,之后有监督微调使模型专门用于双向生成。然而,掩码扩散也引入了自身的错误来源,因为模型在干净的真实上下文上训练,但在推理时遇到自生成且可能错误的上下文,在高度掩码状态下犯的错误会传播到后续步骤。我们进一步引入了生成感知的自校正,在训练期间将模型暴露于其自身的预测,并在早期采样步骤中应用校正传递以修正不可靠地提交的令牌。在HumanML3D和KIT-ML上的大量实验表明,在两个任务上都取得了有竞争力的性能,验证了所提出的两阶段训练和自校正设计的有效性。项目页面可从此https URL访问。
英文摘要
Text-to-motion generation and motion-to-text captioning are two fundamental tasks in human motion modeling, both grounded in the same underlying motion-text correspondence. Existing unified approaches mostly rely on autoregressive modeling, which imposes a fixed generation order and is therefore poorly suited to the bidirectional dependencies between language and motion, allowing early prediction errors to persist as fixed context and degrade both temporal coherence and cross-modal consistency. Masked discrete diffusion, which models sequences through iterative bidirectional prediction, offers a natural remedy. We therefore propose BiMoGen (Bidirectional Motion-text Generation), a unified masked discrete diffusion framework for bidirectional motion-text modeling. To stabilize training, we design Decoupled Uni- and Cross-Modal Training, in which masked pretraining first establishes cross-modal correspondence on paired motion-text sequences, after which supervised fine-tuning specializes the model for bidirectional generation. Masked diffusion nonetheless introduces its own source of error, as the model is trained on clean ground-truth context yet encounters self-generated and potentially erroneous context at inference, with errors committed under heavily masked states propagating through subsequent steps. We further introduce Generation-Aware Self-Correction that exposes the model to its own predictions during training and applies correction passes at early sampling steps to revise unreliably committed tokens. Extensive experiments on HumanML3D and KIT-ML demonstrate competitive performance on both tasks, validating the effectiveness of the proposed two-stage training and self-correction designs. The project page is available at https://wengwanjiang.github.io/BiMoGen-Page.
CommentsAccepted by NeurIPS 2026