arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SWIM:视觉-语言-接地软体全身交互操作

SWIM: Vision-Language-Grounded Soft Whole-Body Interactive Manipulation

Tingcong Liu, Aye Phyu Phyu Aung, Junjie Xiong, Siyi Ma, Bo An, Ke Wu, Senthilnath Jayavelu

arXiv 2609.17035首次发表:更新:

发表机构

Nanyang Technological University; Institute of Advanced Intelligence and Computing, A*STAR; Mohamed bin Zayed University of Artificial Intelligence; National University of Singapore(南洋理工大学; A*STAR先进智能与计算研究所; 穆罕默德·本·扎耶德人工智能大学; 新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SWIM框架通过视觉-语言-动作策略结合扩散头和软体本体感觉,将RGB观察和语言指令映射为驱动序列,在软体机器人上实现高成功率操作。

AI 中文摘要

软体和连续体机器人通过分布式身体变形和接触实现操作,但将语言和视觉情境转化为可执行的全身驱动仍然是一个基本挑战。我们提出SWIM,一个将初始RGB观察和语言指令映射到完整驱动命令序列的框架。其视觉-语言-动作(VLA)策略SWIM-VLA,通过RGB观察、语言指令和肌腱状态的共享表示,将扩散动作头与视觉软体本体感觉(VSP)相结合。扩散头对专家命令块的条件分布进行建模,而VSP利用仿真地面真值监督有序身体锚点预测,鼓励表示在从有限演示学习时保留身体几何结构。具身机械智能支持通过从演化模拟观察中迭代虚拟回滚生成的命令序列的物理执行,内在柔顺性提供局部接触适应,无需在线策略查询。我们在平面肌腱驱动软体机器人上评估SWIM的打包、到达和抓取任务,抓取目标被锚定。在仿真中,SWIM-VLA分别达到100%、96%和88%的成功率,优于改编的OpenVLA-OFT基线和受控消融。在硬件上,SWIM分别达到100%、80%和75%的成功率,而相同策略检查点的直接在线部署分别为75%、40%和25%。

英文摘要

Soft and continuum robots enable manipulation through distributed body deformation and contact, yet translating language and visual context into executable whole-body actuation remains a fundamental challenge. We present SWIM, a framework that maps an initial RGB observation and a language instruction to a complete actuation-command sequence. Its vision-language-action (VLA) policy, SWIM-VLA, combines a diffusion action head with Visual Soft Proprioception (VSP) through a shared representation of RGB observations, language instructions, and tendon states. The diffusion head models conditional distributions of expert command chunks, while VSP supervises ordered body-anchor predictions using simulation ground truth, encouraging the representation to retain body geometry when learning from limited demonstrations. Embodied mechanical intelligence supports physical execution of command sequences generated through iterative virtual rollout from evolving simulated observations, with intrinsic compliance providing local contact adaptation without online policy queries. We evaluate SWIM on packing, reaching, and grasping on a planar tendon-driven soft robot, with grasping targets anchored. In simulation, SWIM-VLA achieves success rates of 100\%, 96\%, and 88\%, respectively, outperforming an adapted OpenVLA-OFT baseline and controlled ablations. On hardware, SWIM achieves success rates of 100\%, 80\%, and 75\%, compared with 75\%, 40\%, and 25\% for direct online deployment of the same policy checkpoint.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑