RAPAC-DP:延迟执行场景下扩散策略的响应对齐待执行动作补偿
RAPAC-DP: Response-Aligned Pending-Action Compensation for Diffusion Policies under Delayed Execution
浏览论文内容
中文总结 AI 辅助
本文针对云端部署模仿学习策略的延迟问题,提出RAPAC-DP框架,通过编码待执行动作补偿延迟,在Kinetix和RoboMimic任务中取得良好性能,验证了该方法的有效性。
中文摘要 AI 辅助
云端推理可为模仿学习策略提供更强大的计算资源,但通信与计算延迟会降低控制性能。为补偿这些延迟,本文提出RAPAC-DP,这是一种响应对齐待执行动作补偿框架,适用于扩散型和流型动作生成器。RAPAC-DP会在云端响应到达前,将已计划执行的动作编码为待执行动作序列,作为参数高效补偿路径的条件输入。当延迟影响可忽略时,绕过该路径可精确恢复冻结的基础策略。训练时,RAPAC-DP从无延迟演示中构建延迟条件样本,无需显式系统动力学或额外延迟演示。在Kinetix数据集测试的最大固定延迟下,RAPAC-DP保留了81.4%的无延迟整体性能;在RoboMimic各任务测试的最大固定延迟下,其在三个任务中的平均成功率达0.633。这些结果证明了待执行动作补偿对云端部署的模仿学习策略的有效性。
英文摘要
Cloud-side inference gives imitation-learning policies access to greater computational resources, but communication and computation delays can degrade control performance. To compensate for these delays, we propose RAPAC-DP, a response-aligned pending-action compensation framework designed for both diffusion- and flow-based action generators. RAPAC-DP encodes the actions already scheduled for execution before the cloud response arrives into a pending-action sequence that serves as the conditioning input to a parameter-efficient compensation pathway. When delay effects are negligible, bypassing this pathway exactly recovers the frozen base policy. For training, RAPAC-DP constructs delay-conditioned samples from delay-free demonstrations, requiring neither explicit system dynamics nor additional delayed demonstrations. At the largest fixed delay tested on Kinetix, RAPAC-DP retained 81.4% of its overall delay-free performance. At the largest fixed delay tested on each RoboMimic task, it achieved a mean success rate of 0.633 across the three tasks. These results demonstrate the effectiveness of pending-action compensation for cloud-deployed imitation-learning policies.
发表机构
- Zhejiang University(浙江大学)
- Midea Group(美的集团)
机构由 AI 辅助整理,请以论文原文为准。