AI 中文总结
本文提出命令-状态差异加权(CSDW)方法,利用命令与状态间的差异信息改进受限交互下的机器人模仿学习,无需阶段注释或架构改动,在真实机器人任务上验证了其有效性。
AI 中文摘要
从测量到的机器人运动构建动作目标是模仿学习中的一种成熟方法。然而,在交互约束下,命令-状态差异可能反映运动本身无法捕捉的控制需求。我们研究了这种信息何时重要以及如何利用它。在三个真实机器人任务中,任务和阶段分析揭示了在受限交互下更大的监督差距,而选择性命令保留提供了局部有用命令信息的证据。基于这些发现,我们提出了命令-状态差异加权(CSDW),该方法考虑机器人响应时间,并将后续进展、持续未满足需求和需求变化组合成连续权重用于命令监督。该方法不需要任务阶段注释,也不改变策略架构或推理。CSDW在受限任务上优于统一命令监督,而在约束较少的任务中,各方法表现相似。
英文摘要
Constructing action targets from measured robot motion is an established approach in imitation learning. Under interaction constraints, however, command-state discrepancy may reflect control demands that motion alone does not capture. We investigate when this information matters and how to exploit it. Across three real-robot tasks, task and phase analyses reveal larger supervision gaps under constrained interaction, while selective command retention provides evidence of locally useful command information. Building on these findings, we propose Command-State Discrepancy Weighting (CSDW), which accounts for robot response times and combines subsequent progress, persistent unmet demand, and demand changes into continuous weights for command supervision. The method requires no task-phase annotations or changes to policy architecture or inference. CSDW improves over uniform command supervision on constrained tasks, while methods perform similarly in the less constrained task. Project page: https://seen-e.github.io/CSDW/.
Comments8 pages, 10 figures. Corrected affiliation name to SEEN-E Robotics, added the project page link to the abstract, and clarified wording and formatting. Methods and experimental results unchanged