UniMoFlow:将指令驱动的3D人体动作编辑建立在生成基础上
UniMoFlow: Grounding Instruction-Driven 3D Human Motion Editing in Generation
浏览论文内容
中文总结 AI 辅助
UniMoFlow将指令驱动3D人体动作编辑建立在文本到动作生成基础上,通过构建Omni-MoEdit数据集、提出UniMoFlow模型与SAFE方法,提升了编辑效果与生成质量
中文摘要 AI 辅助
指令驱动的3D人体动作编辑需要精确的时空定位、丰富的语义接地以及对未修改内容的严格保留。现有方法要么依赖生成模型的无训练适配,要么仅依赖三元组监督;然而,适配常产生次优控制,而人工构建的三元组数据集在规模和语义多样性上仍严重受限。为克服这一瓶颈,我们在数据、架构和推理层面将动作编辑直接建立在文本到动作生成的基础上。在数据层面,我们开发了闭环合成与验证流程,生成了Omni-MoEdit,这是一个涵盖身体部位、幅度、时间、动作和风格编辑的大规模数据集。在架构层面,我们引入了UniMoFlow,这是一个统一的潜在流匹配模型,在生成和编辑之间共享广泛的语义和运动学知识。在推理层面,SAFE(源锚定流编辑)为UniMoFlow补充了可控的、源锚定的优化。此外,我们用语义感知指标增强了标准评估,以考虑那些本质上偏离单一真实参考的有效编辑。大量实验表明,该方法在目标文本对齐、编辑有效性和循环一致性方面有所提升,同时保持了具有竞争力的源保真度和文本到动作生成质量。
英文摘要
Instruction-driven editing of 3D human motion requires precise spatiotemporal localization, rich semantic grounding, and strict preservation of unmodified content. Existing methods either resort to training-free adaptation of generative models or rely solely on triplet supervision; however, adaptation often yields suboptimal control, and manually curated triplet datasets remain severely limited in scale and semantic diversity. To overcome this bottleneck, we ground motion editing directly within text-to-motion generation across data, architecture, and inference. At the data level, we develop a closed-loop synthesis-and-verification pipeline that produces Omni-MoEdit, a large-scale dataset spanning body-part, amplitude, temporal, action, and style edits. At the architectural level, we introduce UniMoFlow, a unified latent flow-matching model that shares broad semantic and kinematic knowledge between generation and editing. At the inference level, SAFE (Source-Anchored Flow Editing) complements UniMoFlow with controllable, source-anchored refinement. Furthermore, we augment standard evaluations with semantics-aware metrics to account for valid edits that inherently deviate from a single ground-truth reference. Extensive experiments demonstrate improved target-text alignment, edit effectiveness, and cycle consistency, while maintaining competitive source fidelity and text-to-motion generation quality.