弥合具身差距:在软体机器人上部署视觉-语言-动作模型
Bridging Embodiment Gaps: Deploying Vision-Language-Action Models on Soft Robots
- EPFL(苏黎世联邦理工学院)
- LatentWorlds AI
- TUDelft(代尔夫特理工大学)
- Embodied AI SA(具身人工智能股份有限公司)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
本文提出结构化微调和部署流程,将 OpenVLA-OFT 与 π₀ 部署到软体连续体机械臂,证明微调可弥合具身差距并实现安全人机交互。
中文摘要 AI 辅助
机器人系统越来越被期望在以人为中心、非结构化的环境中运行,而安全性、适应性和泛化能力在这些环境中至关重要。视觉-语言-动作(VLA)模型已被提出作为一种面向真实机器人的语言引导式通用控制框架。然而,其部署一直局限于传统的串联连杆机械臂。再加上其刚性以及基于学习的控制所具有的不可预测性,安全地与环境交互的能力仍然缺失,却又十分关键。在本工作中,我们展示了在软体连续体机械臂上部署 VLA 模型,以演示自主、安全的人机交互。我们提出了一种结构化的微调和部署流程,在具有代表性的操作任务上评估两个最先进的 VLA 模型(OpenVLA-OFT 和 π₀),并表明:尽管开箱即用策略会因具身形态不匹配而失败,但通过有针对性的微调,软体机器人能够取得与刚性对应物同等的表现。我们的发现强调了微调对于弥合具身差距的必要性,并证明将 VLA 模型与软体机器人相结合,能够在人类共享环境中实现安全、灵活的具身 AI。
英文摘要
Robotic systems are increasingly expected to operate in human-centered, unstructured environments where safety, adaptability, and generalization are essential. Vision-Language-Action (VLA) models have been proposed as a language guided generalized control framework for real robots. However, their deployment has been limited to conventional serial link manipulators. Coupled by their rigidity and unpredictability of learning based control, the ability to safely interact with the environment is missing yet critical. In this work, we present the deployment of a VLA model on a soft continuum manipulator to demonstrate autonomous safe human-robot interaction. We present a structured finetuning and deployment pipeline evaluating two state-of-the-art VLA models (OpenVLA-OFT and $π_0$) across representative manipulation tasks, and show while out-of-the-box policies fail due to embodiment mismatch, through targeted finetuning the soft robot performs equally to the rigid counterpart. Our findings highlight the necessity of finetuning for bridging embodiment gaps, and demonstrate that coupling VLA models with soft robots enables safe and flexible embodied AI in human-shared environments.