AI 中文总结
本文针对形态特殊的缆绳驱动并联机器人,无需实体专属数据集,通过仿真中PPO与GRPO两阶段强化学习,提升了OpenVLA-OFT的指令执行成功率,为仅靠RL生成适配全新实体的语言条件控制器提供了证据。
AI 中文摘要
将预训练的视觉-语言-动作(VLA)策略适配到新机器人通常需要特定于实体的演示,这一假设对于形态与大型机器人数据集中的机械臂差异极大的定制机器人来说尤其受限。本文研究更具挑战性的场景:在无演示的情况下将OpenVLA-OFT适配到带简单夹具、拥有全新控制接口的缆绳驱动并联机器人(CDPR)。本文未采用监督微调,而是利用仿真中的强化学习,基于仿真状态计算密集型几何奖励。训练分两个阶段进行:第一阶段为PPO阶段,用于学习方向运动原语;第二阶段从PPO检查点开始,采用GRPO算法,扩展指令空间以包含物体条件指令。在四个共享方向指令上,PPO后的平均保留成功率为34.25%,经PPO→GRPO后提升至53.50%,其中“向左移动”和“向后移动”的提升尤为显著。GRPO阶段额外引入“移动到<物体>”指令,覆盖8个目标物体,严格成功率为9.75%(39/400),定性 rollout 结果显示,在后期不稳定前常出现正确的目标导向接近行为。与依赖演示数据集且多针对标准刚性臂实体的现有OpenVLA和OpenVLA-OFT结果相比,本文方法完全不使用特定于实体的数据集。结果虽尚未实现稳健操作,但提供了更强证据:仅基于RL引导可生成首个适用于全新实体的可用语言条件控制器。
英文摘要
Adapting a pretrained vision-language-action (VLA) policy to a new robot usually assumes embodiment-specific demonstrations. This assumption is especially restrictive for custom robots whose morphology differs strongly from the manipulators seen in large robot datasets. We study a harder setting: zero-demo embodiment alignment of OpenVLA-OFT on a cable-driven parallel robot (CDPR) with a simple gripper and a previously unseen control interface. Instead of supervised fine-tuning, we use reinforcement learning in simulation with dense geometric rewards computed from simulator state. The training is performed in two stages: a PPO stage for directional motion primitives, followed by GRPO continuation from the PPO checkpoint with an expanded instruction space that includes object-conditioned commands. On the four shared directional instructions, the average held-out success rate improves from 34.25\% after PPO to 53.50\% after PPO$\rightarrow$GRPO, with especially large gains on \texttt{move left} and \texttt{move backward}. In the GRPO stage we additionally introduce \texttt{move to <object>} over eight target objects and obtain 39/400 = 9.75\% strict success, while qualitative rollouts frequently show correct target-directed approach behavior before late-stage instability. Compared with prior OpenVLA and OpenVLA-OFT results, which rely on demonstration datasets and mostly standard rigid-arm embodiments, our method uses no embodiment-specific dataset at all. The results do not yet establish robust manipulation, but they provide stronger evidence that RL-only bootstrapping can create the first usable language-conditioned controller for a genuinely novel embodiment.