arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过软体机器人的分布式全臂交互学习推理与操作

Learning to infer and manipulate through distributed whole-arm interaction in a soft robot

Chuhan Zhang, Ebrahim Shahabi, Kseniia Khomenko, Wei Pan, Cosimo Della Santina

arXiv 2608.30773首次发表:更新:

发表机构

Faculty of Mechanical Engineering; School of Engineering, Newcastle University; Delft University of Technology(机械工程学院; 纽卡斯尔大学工程学院; 代尔夫特理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对软体机器人交互利用不足问题,提出含预训练探索策略等创新的强化学习框架,通过混合刚柔臂的IMU实现盲全臂抓取,成功完成物体识别与抓取。

AI 中文摘要

大象、章鱼等动物借助柔性附肢与环境间丰富的大面积交互,将获取物体的非视觉信息与物理操作融为一体。软体机器人为将这一原理应用于工程系统提供了天然平台,但当前机器人智能对物理交互的利用有限,主要将其视为需抑制的干扰,或至多用于补偿物体错位。本文提出一种物理智能框架,其中分布式柔性交互既能揭示任务相关信息,又能组织操作行为,这形成了固有部分可观测问题:关键任务相关信息无法直接测量,必须从物理交互历史中推断。我们提出一种强化学习架构,通过端到端学习基于记忆的控制策略解决该挑战,核心创新包括:(i)预训练的探索策略,为广工作空间探索提供参考;(ii)联合优化,在单个循环策略中整合探索与抓取目标;(iii)两阶段仿真到真实域适配,包括观测映射与策略微调。我们通过混合刚柔机械臂的盲全臂抓取验证该原理,该机械臂的柔性结构中嵌入了IMU,作为本体感知的唯一来源。学习到的策略通过自主协调工作空间探索、物体接触与定位、抓取相关属性的推理以及稳定的全臂包裹,成功识别并抓取各类物体。

英文摘要

In animals such as elephants and octopuses, acquiring non-visual information about an object and physically engaging with it are inseparable processes mediated by rich, large-area interactions between compliant appendages and the environment. Soft robots provide a natural platform for translating this principle into engineered systems. Yet current robotic intelligence makes limited use of physical interaction, treating it primarily as a disturbance to be rejected or, at best, as a means of compensating for object misalignment. Here, we introduce a physical intelligence framework in which distributed compliant interactions jointly reveal task-relevant information and organize manipulation behavior. This results in an intrinsically partially observable problem: key task-relevant information is never measured directly, but must instead be inferred from the history of physical interactions. We propose a reinforcement-learning architecture that addresses this challenge by learning a memory-based control policy end-to-end. The key innovations making this possible are (i) a pretrained exploration policy that provides a reference for broad workspace exploration, (ii) joint optimization that integrates exploration and grasping objectives within a single recurrent policy, and (iii) a two-stage sim-to-real adaptation including observation mapping and policy fine-tuning. We demonstrate this principle through blind whole-arm grasping with a hybrid rigid-soft robotic arm that we equip with IMUs embedded directly within its compliant structure, providing its only source of proprioceptive sensing. The learned policy successfully identifies and grasps various objects by autonomously coordinating workspace exploration, object encounter and localization, inference of grasp-relevant properties, and stable whole-arm wrapping.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑