SuperMotion:文本驱动人体运动编辑的源保持去噪框架
SuperMotion: Source-Preserving Denoising for Text-Driven Human Motion Editing
- King Abdullah University of Science and Technology(阿卜杜拉国王科技大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对文本驱动人体运动编辑中源内容保留不足的问题,提出SuperMotion框架,通过显式重用源运动并学习保留门,在扩散去噪中注入时间细节,提升编辑精度至33.20% R@1。
AI中文摘要:
文本驱动的人体运动编辑旨在实现请求的更改,同时保留兼容的源内容。现有的扩散编辑方法在很大程度上依赖于学习到的条件约束来保留未编辑部分,然而随着去噪过程的进行,其输出可能会丢失时间细节。我们提出了源保持去噪框架(SuperMotion),该框架在每个反向步骤中显式地重用源以实现源保留。我们首先将源运动与输出时间线对齐,并预测一个保留门,该门控制跨帧和特征维度的重用。随后,一个干净空间源锚点利用学习到的保留门将预测的干净运动与对齐的源混合,并将修正后的估计直接传递给采样后验。由于对齐的源是实际运动而非回归输出,该锚点注入了样本级的时间细节,而重建训练的去噪器往往会平滑掉这些细节。为了学习有效的源重用,我们根据编辑目标监督锚定估计,并通过时间高频损失匹配其第二时间差分。这些目标不需要显式的编辑掩码。大量实验表明,SuperMotion提高了编辑精度,在MotionFix上达到了33.20%的全池R@1,同时减少了时间细节衰减,并在实现请求更改时保持了运动动态。消融研究证实,学习到的保留门是性能提升的原因,并且它重用源以正确保留未编辑内容。
英文摘要:
Text-driven human motion editing aims to realize a requested change while preserving compatible source content. Existing diffusion editors rely largely on learned conditioning for preservation of the unedited part, yet their outputs can lose temporal detail as denoising proceeds. We propose the \textbf{Source-Preserving Denoising framework (SuperMotion)}, which explicitly reuses the source at each reverse step for source preservation. We first align the source motion to the output timeline and predict a preservation gate that controls reuse across frames and feature dimensions. A clean-space source anchor then utilizes the learned preservation gate to blend the predicted clean motion with the aligned source and passes the corrected estimate directly to the sampling posterior. Because the aligned source is a realized motion rather than a regression output, the anchor injects sample-level temporal detail that a reconstruction-trained denoiser tends to smooth away. To learn effective source reuse, we supervise the anchored estimate against the editing target and match its second temporal differences through a temporal high-frequency loss. These objectives require no explicit edit masks. Extensive experiments show that SuperMotion improves editing accuracy, reaching 33.20\% full-pool R@1 on MotionFix, while reducing temporal-detail attenuation and preserving motion dynamics as it realizes the requested changes. Ablations confirm that the learned preservation gate is responsible for the gain and that it reuses the source to retain the unedited content properly.