arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2506.24086cs.CVcs.CL

MotionGPT3:将人体动作作为第二模态

MotionGPT3: Human Motion as a Second Modality

  • Zhejiang University(浙江大学)
  • Fudan University(复旦大学)
  • ByteDance(字节跳动)
  • The Chinese University of HongKong(香港中文大学)

机构由 AI 辅助整理,请以论文原文为准。

Bingfan Zhu, Biao Jiang, Sunyi Wang, Shixiang Tang, Tao Chen, Linjie Luo, Youyi Zheng, Xin Chen

更新

AI总结:

提出双模态动作-语言模型MotionGPT3,通过双流Transformer与连续VAE编码避免量化误差并减少跨模态干扰,采用三阶段训练策略,在保持SOTA性能的同时显著加速收敛。

AI中文摘要:

随着大型语言模型(LLMs)的快速发展,统一理解与生成的多模态框架变得极具前景,但随着模态和任务数量的增加,其复杂性也日益增加。我们观察到,动作量化引入的近似误差限制了动作质量,并且在单流主干中统一离散文本和连续动作会加剧跨模态干扰。受近期分离不同模态信号的多分支Transformer设计的启发,我们提出了MotionGPT3,这是一种用于理解和生成的双模态动作-语言模型。MotionGPT3使用变分自编码器(VAE)将原始动作编码到连续潜空间中,从而避免了量化引起的伪影,同时利用了预训练语言模型的语义先验。具有共享注意力的双流Transformer保留了特定模态的路径,同时实现了可控的、双向的信息流,这减少了干扰,稳定了优化,并在不降低保真度的情况下经验性地加速了收敛。对于多模态联合训练,先生成后对齐的三阶段调度进一步提高了稳定性并限制了跨任务干扰。实验表明,MotionGPT3在训练损失上实现了2倍更快的收敛,在验证中实现了高达4倍更快的收敛,同时在标准动作理解和动作生成基准上保持了最先进的性能。

英文摘要:

With the rapid progress of large language models (LLMs), multimodal frameworks that unify understanding and generation have become promising, yet they face increasing complexity as the number of modalities and tasks grows. We observe that motion quantization introduces approximation errors that cap motion quality, and that unifying discrete text and continuous motion within a single-stream backbone amplifies cross-modal interference. Motivated by recent multi-branch Transformer designs that separate signals from different modalities, we propose MotionGPT3, a bimodal motion-language model for both understanding and generation. MotionGPT3 encodes raw motion into a continuous latent space using a variational autoencoder (VAE), thereby avoiding quantization-induced artifacts, while leveraging the semantic prior of pretrained language models. A dual-stream Transformer with shared attention preserves modality-specific routes while enabling controlled, bidirectional information flow, which reduces interference, stabilizing optimization, and empirically accelerates convergence without degrading fidelity. For multimodal joint training, a generate-then-align three-stage schedule further improves stability and limits cross-task interference. Experiments show that MotionGPT3 achieves 2x faster convergence in training loss and up to 4x faster convergence in validation, while maintaining state-of-the-art performance on standard motion understanding and motion generation benchmarks.

补充信息

↑