arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21817cs.RO

一种用于基于块的可变长度动作(VLA)操作策略训练与部署的仿真到现实集成流程

A Sim-to-Real Integration Pipeline for Training and Deployment of Chunk-Based VLA Manipulation Policies

  • Institute of Intelligent Systems and Robotics (ISIR), CNRS, Sorbonne University(智能系统与机器人研究所(ISIR),法国国家科学研究中心,索邦大学)

机构由 AI 辅助整理,请以论文原文为准。

Mathilde Kappel, Clémence Grislain, Mohamed Chetouani, Olivier Sigaud, Louis Annabi, Fa\"ız Ben Amar, Stéphane Doncieux, Mahdi Khoramshahi

AI总结:

提出一种开源仿真到现实集成流程,通过仿真轨迹开环重放收集真实数据并闭环评估,解决VLA操作策略训练数据采集成本高的问题,同时提供仿真到现实差距的直接测量。

AI中文摘要:

视觉-语言-动作(VLA)模型已成为将多模态输入(包括语义指令、场景视觉观测和本体感觉观测)映射到机器人动作的主流范式。大多数最先进的模型在末端执行器位姿空间中预测动作,并将其表示为动作块序列。训练和评估这些模型需要大规模收集真实世界演示数据,将机器人动作与相应的视觉和本体感觉观测配对。在真实硬件上收集此类数据通常依赖人类远程操作,这使得过程成本高昂、耗时且难以扩展。我们提出了一种开源的仿真到现实实验协议,以解决这一瓶颈:在仿真中生成的专家轨迹在真实的Franka FR3平台上以开环方式重放,同时记录相应的真实视觉和本体感觉观测,并将其转换为与VLA训练兼容的格式。随后,相同的部署栈以闭环方式复用于在该平台上评估训练好的策略,从而确保数据收集和评估共享相同的硬件配置。由于每次真实记录都与产生它的仿真轨迹配对,该协议还能直接测量仿真到现实的差距。我们在Hugging Face上发布了收集的数据集以及流程源代码,链接为https://this https URL。

英文摘要:

Vision-Language-Action (VLA) models have become a prominent paradigm for mapping multimodal inputs, including semantic instructions, visual observations of the scene, and proprioceptive observations, to robot actions. Most state-of-the-art models predict actions in the end-effector pose space as sequences of action chunks. Training and evaluating these models requires large-scale collections of real-world demonstrations, pairing robot actions with the corresponding visual and proprioceptive observations. Collecting such data on real hardware typically relies on human teleoperation, making the process costly, time-consuming, and difficult to scale. We present an open-source sim-to-real experimental protocol that addresses this bottleneck: expert trajectories generated in simulation are replayed open-loop on a real Franka FR3 setup, where the corresponding real visual and proprioceptive observations are recorded and converted into a format compatible with VLA training. The same deployment stack is then reused, in closed-loop, to evaluate a trained policy on that setup, so that data collection and evaluation share an identical hardware configuration. Because each real recording is paired with the simulated trajectory that produced it, the protocol also yields a direct measurement of the sim-to-real gap. We release the collected datasets on Hugging Face together with the pipeline source code https://gitlab.isir.upmc.fr/kappel/sim2real_public_chunk_control.

补充信息

↑