发表机构
Agency for Defense Development(国防发展局)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究提出扩散重滚动框架用于机器人序列预测,能选择性重新去噪局部稳定区域,实现跨视野迭代修正。通过与其他方法对比评估,在多任务基准测试中取得更好性能,支持结构化重新去噪是可修正机器人序列生成的有效机制。
AI 中文摘要
我们提出了扩散重滚动(Diffusion ReRoll),这是一种基于扩散的机器人序列预测框架,能够在不同视野上进行可修正去噪。现有的基于扩散的序列预测器通常执行单一的单调去噪过程。相比之下,扩散重滚动会选择性地对已局部稳定的区域重新去噪,而其余区域继续去噪,这样重新去噪的区域可利用其他视野的上下文再次细化。这种结构化的重新去噪实现了迭代的跨视野修正,使前后段能够相互修正,同时保持局部一致性。我们基于扩散强制,在长期规划、策略学习和统一视频动作建模等方面,将扩散重滚动与全序列扩散和因果去噪进行评估对比。在OGBench PointMaze和AntMaze上,扩散重滚动在基于匹配引导的规划中,相对于扩散强制平均成功率有21%的相对提升,在匹配目标修复中相对于Diffuser有23%的提升。在扩散策略式动作预测中,在LIBERO - 10多任务基准上,扩散重滚动在不同预测视野和历史长度下相对于扩散策略平均成功率提高了56.5%。在统一视频动作预测中,扩散重滚动提升了策略和逆动力学性能,特别是在分布外评估中,并实现了最佳的动作 - 视频一致性。这些结果支持结构化重新去噪作为可修正机器人序列生成的有效机制。
英文摘要
We propose Diffusion ReRoll, a diffusion-based framework for robotic sequential prediction that enables revisable denoising over horizons. Existing diffusion-based sequence predictors typically perform a single monotonic denoising process. In contrast, Diffusion ReRoll selectively re-noises regions that have become locally stable while the remaining regions continue denoising, so the re-noised regions can be refined again using context from the rest of the horizon. This structured re-noising enables iterative cross-horizon revision, allowing earlier and later segments to revise one another, while maintaining local consistency. We evaluate Diffusion ReRoll against full-sequence diffusion and causal denoising based on Diffusion Forcing across long-horizon planning, policy learning, and unified video-action modeling. On OGBench PointMaze and AntMaze, Diffusion ReRoll achieves relative gains in average success rate of 21% over Diffusion Forcing in matched guidance-based planning and 23% over Diffuser in matched goal-inpainting. In diffusion-policy-style action prediction, Diffusion ReRoll improves average success by 56.5% relative to Diffusion Policy across different prediction horizons and history lengths on the LIBERO-10 multi-task benchmark. In unified video-action prediction, Diffusion ReRoll improves policy and inverse dynamics performance, especially under out-of-distribution evaluation, and achieves the best action-video consistency. These results support structured re-noising as an effective mechanism for revisable robotic sequence generation.
CommentsProject Page: https://seonsoo-p1.github.io/DiffusionReRoll/