发表机构
The University of Tokyo; Japan Advanced Institute of Science and Technology; Shanda Group(东京大学; 日本先端科学技术大学院大学; 盛大集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出三角重采样(TR)训练后方法,通过扩展回放训练和真实值钳制,缓解运动扩散模型长时程误差累积,在120秒生成任务上实现最先进FID AUC。
AI 中文摘要
我们提出了三角重采样(Triangular Resampling, TR),一种用于缓解运动扩散模型中长时程误差累积的训练后方法。该方法基于FloodDiffusion的三角去噪调度,解决了基于真实值生成的训练窗口与模型生成的推理状态之间的不匹配问题。仅替换已完成的运动历史无法解决活动窗口内部分去噪状态中的这种不匹配。因此,TR将基于回放的训练扩展到这些状态,并利用真实值钳制来限制过度的漂移。对于每个重放样本,TR抽取一个去噪阈值,该阈值在潜在位置和重放更新之间共享,并在不进行梯度跟踪的情况下重放多步三角去噪。每次更新后,低于阈值的状态被替换为与噪声匹配的真实值,而高于或等于阈值的状态保留模型预测。由此产生的潜在窗口进入标准训练更新。这种回放构建支持监督训练(TR)和分布匹配(TR-DMD)。在基于HumanML3D测试提示的120秒运动生成中,TR和TR-DMD分别在其非DMD和DMD比较组中达到了最先进的FID AUC。相对于无回放的匹配训练后方法,监督式TR将FID AUC降低了40.9%,将FID退化斜率降低了55.3%。
英文摘要
We introduce Triangular Resampling (TR), a post-training method for mitigating long-horizon error accumulation in motion diffusion models. Built on FloodDiffusion's triangular denoising schedule, TR addresses the mismatch between ground-truth-derived training windows and model-generated inference states. Replacing only completed motion history leaves this mismatch unresolved in partially denoised states within the active window. TR therefore extends rollout-based training to these states, using ground-truth clamping to limit excessive drift. For each replayed sample, TR draws one denoising threshold, shared across latent positions and replay updates, and replays multi-step triangular denoising without gradient tracking. After each update, states below the threshold are replaced with noise-matched ground truth, while those at or above it retain model predictions. The resulting latent window enters the standard training update. This rollout construction supports both supervised training (TR) and distribution matching (TR-DMD). On 120-second motion generation from HumanML3D test prompts, TR and TR-DMD achieve state-of-the-art FID AUC within their respective non-DMD and DMD comparison groups. Supervised TR reduces FID AUC by 40.9% and FID degradation slope by 55.3% relative to matched post-training without replay.