SeedPolicy: 通过自进化扩散策略实现机器人操控的水平扩展
SeedPolicy: Horizon Scaling via Self-Evolving Diffusion Policy for Robot Manipulation
浏览论文内容
中文总结 AI 辅助
本文提出SEGA模块,通过门控注意力机制提升扩散策略在长时域操控中的表现,实验表明其在RoboTwin 2.0基准上优于现有方法,具有更高的效率和性能。
中文摘要 AI 辅助
模仿学习(IL)使机器人能够从专家示范中获取操控技能。扩散策略(DP)能建模多模态专家行为,但当简单增加堆叠观测时域时会退化,限制了长时域操控。我们提出自进化门控注意力(SEGA),一种时间模块,通过门控注意力维护时间演化的潜在状态,实现高效的递归更新,将长期上下文累积到紧凑的潜在表示中,同时过滤无关的时序信息。将SEGA集成到DP中得到自进化扩散策略(SeedPolicy),解决了时间建模瓶颈,通过适度开销扩展有效的时间时域。在具有50个操控任务的RoboTwin 2.0基准上,SeedPolicy优于DP和其他IL基线。在CNN和Transformer后端平均情况下,SeedPolicy在干净设置中实现36.8%的相对改进,在随机挑战设置中实现169%的相对改进。与参数达12亿的视觉-语言-动作模型如RDT相比,SeedPolicy在干净设置中以数量级更少的参数实现了更强的性能,证明了其强大的效率。这些结果确立了SeedPolicy作为长时域机器人操控模仿学习的最先进方法。代码可在:https://anonymous.4open.science/r/SeedPolicy-64F0/获得。
英文摘要
Imitation Learning (IL) enables robots to acquire manipulation skills from expert demonstrations. Diffusion Policy (DP) models multi-modal expert behaviors but degrades when naively increasing stacked observation horizons, limiting long-horizon manipulation. We propose Self-Evolving Gated Attention (SEGA), a temporal module that maintains a time-evolving latent state via gated attention, enabling efficient recurrent updates that accumulate long-term context into a compact latent representation while filtering irrelevant temporal information. Integrating SEGA into DP yields Self-Evolving Diffusion Policy (SeedPolicy), which resolves the temporal modeling bottleneck and extends the effective temporal horizon with moderate overhead. On the RoboTwin 2.0 benchmark with 50 manipulation tasks, SeedPolicy outperforms DP and other IL baselines. Averaged across both CNN and Transformer backbones, SeedPolicy achieves 36.8% relative improvement in clean settings and 169% relative improvement in randomized challenging settings over the DP. Compared to vision-language-action models such as RDT with 1.2B parameters, SeedPolicy achieves stronger performance in the clean setting with one to two orders of magnitude fewer parameters, demonstrating strong efficiency. These results establish SeedPolicy as a state-of-the-art imitation learning method for long-horizon robotic manipulation. Code is available at: https://github.com/Youqiang-Gui/SeedPolicy.
发表机构
- Sichuan University(四川大学)
- Dexmal Inc.(Dexmal公司)
- Independent Researcher(独立研究员)
- UESTC
机构由 AI 辅助整理,请以论文原文为准。