arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34018cs.ROcs.LG

估计而非模仿:复用可微的基于状态策略进行视觉运动控制

Estimate, Don't Imitate: Reusing Differentiable State-Based Policies for Visuomotor Control

  • Fraunhofer Heinrich-Hertz-Institut(弗劳恩霍夫海因里希·赫兹研究所)
  • Technische Universität Berlin(柏林工业大学)
  • BIFOLD – Berlin Institute for the Foundations of Learning and Data(BIFOLD – 柏林学习与数据基础研究所)
  • Robotics Institute Germany(德国机器人研究所)

机构由 AI 辅助整理,请以论文原文为准。

Denis Shcherba, Adrian Abel, Eckart Cobo-Briesewitz, Paul Mattes, Wojciech Samek, Marc Toussaint

AI总结:

本文提出复用可微的基于状态专家策略,仅训练视觉状态估计器并加入动作一致性损失,在五个操作任务中优于直接模仿,并实现76%的仿真到现实迁移成功率。

AI中文摘要:

仿真训练的操作策略可以利用特权状态信息学习有效的接触丰富行为,但部署时需要从部分观测(如噪声相机图像)中行动。一种常见解决方案是师生蒸馏,其中训练视觉运动策略以复现特权专家的动作。这要求学生联合推断任务相关状态并重新学习专家已有的动作映射。另一种方案是复用基于状态的专家,仅学习一个感知接口来重建其缺失的状态输入。然而,仅最小化状态估计误差并不一定能最小化由这些估计引起的下游控制误差。为弥合这一差距,我们使用直接状态监督和通过冻结的、可微专家反向传播的动作一致性损失来训练视觉状态估计器。一个调度目标首先建立物理上有意义的状态估计,并逐步强调影响专家动作的误差。在五个目标条件操作任务中,保留专家始终优于从相同专家演示语料库进行直接像素到动作的模仿。我们进一步在物理Panda机器人上展示了仿真到现实的迁移,在无需重新训练底层专家的情况下实现了76%的成功率。

英文摘要:

Simulation-trained manipulation policies can exploit privileged state information to learn effective contact-rich behaviours, but deployment requires acting from partial observations such as noisy camera images. A common solution is teacher-student distillation, in which a visuomotor policy is trained to reproduce the actions of the privileged expert. This requires the student to jointly infer the task-relevant state and relearn the expert's action mapping that is already available. An alternative is to reuse the state-based expert and learn only a perceptual interface that reconstructs its missing state inputs. However, minimising the state estimate error alone does not necessarily minimise the downstream control error induced by these estimates. To bridge this gap, we train a visual state estimator using both direct state supervision and an action-consistency loss backpropagated through the frozen, differentiable expert. A scheduled objective first establishes a physically meaningful state estimate and progressively emphasises errors that affect the expert's actions. Across five goal-conditioned manipulation tasks, retaining the expert consistently outperforms direct pixel-to-action imitation from the same expert demonstration corpus. We further demonstrate sim-to-real transfer on a physical Panda robot, achieving 76% success without retraining the underlying expert.

补充信息

↑