发表机构
East China Normal University; School of Computer Science and Technology, East China Normal University; School of Computer Science and Engineering, Shanghai Jiao Tong University; Shanghai Institute of Artificial Intelligence for Education, East China Normal University; School of Statistics, East China Normal University; School of Artificial Intelligence, Beihang University(华东师范大学; 华东师范大学计算机科学与技术学院; 上海交通大学计算机科学与工程学院; 华东师范大学上海教育人工智能研究院; 华东师范大学统计学院; 北京航空航天大学人工智能学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对文本驱动3D人体动作编辑中变化与不变性难以平衡的问题,提出CIME框架,通过空间姿态和时间节奏维度的解耦实现最优性能,相关成果已开源。
AI 中文摘要
文本驱动的人体动作编辑旨在根据自然语言指令修改现有动作序列,同时保持原始动作的结构一致性。现有的基于扩散的方法难以平衡对文本做出响应的“变化”与惯性带来的“不变性”,它们通常依赖粗糙的空间约束和刚性的均匀时间假设,导致在可变长度编辑过程中出现空间动作扭曲和固有物理节奏被破坏的问题。为应对这些挑战,我们提出了变化与不变性动作编辑(Change and Invariance Motion Editing,CIME),这是一个将变化与不变性全面解耦为空间姿态和时间节奏维度的统一框架。对于空间姿态,我们的方法整合了包含分层回溯特征监督、细微动作保留和基于三元组的语义对齐的全监督正负学习机制;对于时间节奏,我们引入了黎曼非均匀积分流形映射(Riemannian Non-uniform Integral Manifold Mapping,RNIMM)模块,该模块通过运动学感知的非均匀时间戳实现编辑文本中物理节拍的高保真复现。在MotionFix和STANCE Adjustment数据集上开展的大量实验表明,CIME在编辑对齐度和结构保真度方面达到了最先进的性能,验证了我们统一架构的有效性。我们的源代码和模型已发布在此http URL。
英文摘要
Text-driven human motion editing aims to modify existing motion sequences according to natural language instructions while maintaining the structural consistency of the original motion. Existing diffusion-based approaches struggle to balance text-responsive "change" and inertial "invariance". They often rely on coarse spatial constraints and rigid uniform time assumptions, leading to spatial motion distortions and the destruction of intrinsic physical rhythms during variable-length editing. To handle these challenges, we propose Change and Invariance Motion Editing (CIME), a unified framework that comprehensively decouples change and invariance into spatial pose and temporal rhythm dimensions. For spatial poses, our method integrates an omni-supervised positive-negative learning mechanism comprising hierarchical retrospective feature supervision, subtle motion preservation, and triplet-based semantic alignment. For temporal rhythms, we introduce the Riemannian Non-uniform Integral Manifold Mapping (RNIMM) module, which achieves high-fidelity reproduction of physical beats in the edited text via kinematics-aware non-uniform timestamps. Extensive experiments on the MotionFix and STANCE Adjustment datasets demonstrate that CIME achieves state-of-the-art performance in editing alignment and structural fidelity, validating the effectiveness of our unified architecture. Our source codes and models have been released at: github.com/ZhenwuShi/CIME.git