FloodDiffusion 2:高效且路径可控的流式运动生成
FloodDiffusion 2: Efficient and Path Controllable Streaming Motion Generation
浏览论文内容
中文总结 AI 辅助
FloodDiffusion 2通过部分注意力、Bregman准则和路径条件控制,实现高效且可控的流式运动生成,训练计算减少4.6倍,去噪加速11.29倍,FID在SEED和HumanML3D上分别达0.048和0.053。
中文摘要 AI 辅助
我们提出了FloodDiffusion 2(FD2),一个高效且可控的框架,它建立在FloodDiffusion(FD1)——一种最先进的流式运动生成模型——的基础上。虽然FD1能生成合理的运动,但它存在效率低和可控性有限的问题,因为其注意力设计需要对整个历史进行重复计算,并且缺乏面向实际应用的精确轨迹控制。为解决这些局限并提升生成质量,FD2引入了三项改进。首先,部分注意力(Partial Attention)使最终化的历史表示独立于活动窗口,从而支持KV缓存推理和共享历史打包,实现高效训练。其次,我们为回归损失建立了一个必要且充分的Bregman准则,以保持扩散的条件均值速度场。该准则指导了一种由正向运动学(FK)诱导的二次损失,在不进行在线FK评估的情况下融入运动几何。第三,FD2引入了精确的路径条件控制,以控制角色的根轨迹,同时保持自然的身体运动。实验表明,FD2将训练计算量减少了4.6倍,并将去噪速度加快了11.29倍,在长序列上达到每次更新2.303毫秒。除了这些效率提升,FD2还提高了运动质量,超越了FD1,并在流式方法中取得了最先进的FID分数,在SEED上为0.048,在HumanML3D上为0.053。
英文摘要
We present FloodDiffusion 2 (FD2), an efficient and controllable framework that builds upon FloodDiffusion (FD1), a state-of-the-art streaming motion generation model. While FD1 produces plausible motion, it suffers from low efficiency and limited controllability, as its attention design requires repeated computation over the entire history, and it lacks precise trajectory control for real-world applications. To address these limitations and improve generation quality, FD2 introduces three advances. First, Partial Attention makes finalized history representations independent of the active window, enabling KV-cached inference and shared-history packing for efficient training. Second, we establish a necessary-and-sufficient Bregman criterion for regression losses to preserve diffusion's conditional-mean velocity field. This criterion guides an FK-induced quadratic loss that incorporates motion geometry without online FK evaluation. Third, FD2 introduces precise path conditioning to control the character's root trajectory while preserving natural body motion. Experiments show that FD2 reduces training computation by 4.6$\times$ and accelerates denoising by 11.29$\times$, reaching 2.303 ms per update on long sequences. Alongside these efficiency gains, FD2 improves motion quality over FD1 and achieves state-of-the-art FID scores among streaming methods, with 0.048 on SEED and 0.053 on HumanML3D.