arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.40134cs.ROcs.AI

触觉好奇心驱动机器人交互

Tactile Curiosity Drives Robot Interaction

  • ETH Zürich(苏黎世联邦理工学院)
  • University of California, Berkeley(加利福尼亚大学伯克利分校)
  • The University of Texas at Austin(得克萨斯大学奥斯汀分校)

机构由 AI 辅助整理,请以论文原文为准。

Klemens Iten, Alexander Proshkin, Bhavya Sukhija, Stelian Coros, Andreas Krause, Pieter Abbeel, Carmelo Sferrazza

AI总结:

本文提出TacEx框架,利用触觉反馈引导好奇心驱动的探索,使机器人无需任务奖励即可学习操作技能,并提升下游策略及VLA模型性能。

AI中文摘要:

通过强化学习(RL)掌握机器人操作技能在很大程度上仍然样本效率低下。最常见的RL算法依赖于随机动作采样来发现新策略,导致智能体将大部分训练预算用于自由空间中的运动,远离了操作技能得以产生的接触。现有的基于模型分歧或认知不确定性的内在动机方法改善了各向同性噪声,但它们也可能奖励功能无关转换中的不确定性,例如自由空间中的不稳定运动。在这项工作中,我们认为触觉反馈为探索提供了自然的信号,并引入了TacEx框架,该框架通过跨感官模态分解模型不确定性并将好奇心导向触觉通道,将触摸整合到认知不确定性驱动的探索中。通过将好奇心锚定于触觉,TacEx驱动机器人发现复杂的接触动力学,在探索过程中无需任务奖励或专家演示即可学习操作和抓取物体。通过这种触觉驱动的探索收集的交互密集数据集支持下游拾取和放置策略的离线学习,无需额外的环境交互。我们进一步使用触觉驱动的探索来对视觉-语言-动作(VLA)模型进行后训练。尽管VLA最初在没有触觉反馈的情况下进行预训练,但使用TacEx进行后训练显著提高了下游性能,同时保持了高样本效率。

英文摘要:

Mastering robot manipulation skills via reinforcement learning (RL) remains largely sample-inefficient. The most common RL algorithms rely on random action sampling to discover new strategies, resulting in agents that allocate most of their training budget to motions in free space, away from the contacts from which manipulation skills emerge. Existing intrinsic motivation methods based on model disagreement or epistemic uncertainty improve on isotropic noise, but they can also reward uncertainty in functionally irrelevant transitions, such as erratic motions in free space. In this work, we argue that tactile feedback provides a natural signal for exploration, and introduce TacEx, a framework that incorporates touch into epistemic uncertainty-driven exploration by decomposing model uncertainty across sensory modalities and directing curiosity toward the tactile channel. By anchoring curiosity to the sense of touch, TacEx drives the robot to discover complex contact dynamics, learning to manipulate and grasp objects without task rewards or expert demonstrations during exploration. The interaction-dense dataset collected through this tactile-driven curiosity supports offline learning of downstream pick-and-place policies without additional environment interaction. We further use tactile-driven exploration to post-train vision-language-action (VLA) models. Although the VLAs are initially pre-trained without tactile feedback, post-training with TacEx substantially improves downstream performance while remaining highly sample-efficient.

补充信息

↑