发表机构
Institute of Trustworthy Embodied AI, Fudan University, China; Shanghai Key Laboratory of Multimodal Embodied AI, China(可信具身人工智能研究院,复旦大学,中国; 上海多模态具身人工智能重点实验室,中国)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对机器人模仿学习中现有方法的不足,提出SegDiff闭环视觉运动策略,分解示范为运动段并预测连续轨迹,利用扩散模型等提出动态时间集成机制,在多场景中性能优于现有方法,能处理长时间依赖并保持实时适应与控制稳定。
AI 中文摘要
模仿学习使机器人能通过将观察映射到动作从示范中获取操作技能。现有方法要么预测短期连续动作序列,要么预测离散关键姿态。连续预测方法因预测范围短存在复合误差且难以处理多模态动作分布,基于关键姿态的方法需要外部规划器,限制实时适用性。为应对这些挑战,我们引入SegDiff,一种整合两种范式优势的闭环视觉运动策略。SegDiff将示范分解为关键姿态间的运动段,并学习预测从当前状态到下一个关键姿态的连续轨迹,实现长时预测与实时优化。此外,我们利用扩散模型和DDIM反演能力提出动态时间集成机制,使策略能有效应对动态环境并减轻多模态采样不一致导致的不连续性。SegDiff在各种模拟和现实场景中比现有方法有显著性能提升,表明其在保持实时适应性和控制稳定性的同时,能很好地处理长时间依赖关系。
英文摘要
Imitation learning enables robots to acquire manipulation skills from demonstrations by mapping observations to actions. Existing approaches predict either short-horizon continuous action sequences or discrete keyposes. However, continuous prediction methods suffer from compounding errors due to short prediction horizons and struggle with multi-modal action distributions, whereas keypose-based methods necessitate an external planner, constraining real-time applicability. To address these challenges, we introduce SegDiff, a closed-loop visuomotor policy that integrates the strengths of both paradigms. SegDiff decomposes demonstrations into motion segments between keyposes and learns to predict the continuous trajectory from the current state to the next keypose, enabling long-horizon prediction with real-time refinement. Furthermore, we leverage the capability of diffusion models and DDIM inversion to propose a Dynamic Temporal Ensembling mechanism, which allows the policy to efficiently respond to dynamic environments and mitigate discontinuities caused by inconsistent multi-modal sampling. SegDiff demonstrates significant performance gains over existing approaches across various simulated and real-world scenarios, indicating its strong ability to reason over extended temporal dependencies while maintaining real-time adaptability and control stability.