arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21650cs.ROcs.AI

SynthDemo-RL:利用LLM引导的合成演示突破VLA自适应中的零奖励障碍

SynthDemo-RL: Breaking the Zero-Reward Barrier in VLA Adaptation with LLM-Guided Synthetic Demonstrations

  • Fujitsu Limited(富士通株式会社)
  • The Institute of Statistical Mathematics(统计数理研究所)

机构由 AI 辅助整理,请以论文原文为准。

Hiroaki Kingetsu, Hiroaki Kurihara, Kaoru Yokoo, Kenji Fukumizu, Manohar Kaul

AI总结:

SynthDemo-RL通过LLM引导的合成演示和PPO精炼,解决VLA微调中稀疏奖励的探索难题,在LIBERO-PRO上挽救全部零奖励任务,平均成功率超97%。

AI中文摘要:

微调视觉-语言-动作(VLA)模型通常依赖于人类遥操作演示,而具有稀疏二元奖励的强化学习(RL)在成功轨迹很少被采样时面临探索挑战。我们提出SynthDemo-RL,一个教师-学生框架,其中自动化教师将模拟器特权状态转换为成功的操作轨迹,VLA学生通过监督微调(SFT)从这些轨迹中蒸馏,并使用具有二元任务成功奖励的PPO来精炼学生。我们研究奖励覆盖率,即在固定评估协议下至少观察到一次成功的任务比例,作为平均成功率的补充。在LIBERO-PRO(一个无演示的扰动LIBERO任务公共基准)上,57个评分任务中有27个对于在原始LIBERO任务上微调的pi_0.5策略恰好为0%成功率。从该策略直接进行PPO,使用与SynthDemo-RL精炼阶段相同的PPO配方和相同的RL计算,挽救了这27个任务中的10个,留下17个为0%。SynthDemo-RL,每个任务使用50条合成轨迹且无新的人类演示,挽救了全部27个任务,并在LIBERO-PRO的Position和Task轴上分别达到97.8%和97.1%的平均成功率。在标准LIBERO上,相同的流程在无人类演示的情况下达到96.0%,与使用每个任务50个人类演示训练的pi_0.5相差1.7个百分点。我们进一步在RoboTwin 2.0上验证该流程,并确认在MuJoCo孪生环境中训练的策略生成的轨迹可在物理机器人上开环执行。

英文摘要:

Fine-tuning Vision-Language-Action (VLA) models commonly relies on human teleoperation demonstrations, while reinforcement learning (RL) with sparse binary rewards faces an exploration challenge when successful trajectories are rarely sampled. We propose SynthDemo-RL, a teacher-student framework in which an automated teacher converts simulator-privileged state into successful manipulation trajectories, a VLA student is distilled from them by supervised fine-tuning (SFT), and PPO with binary task-success rewards refines the student. We study reward coverage, the fraction of tasks for which at least one success is observed under the fixed evaluation protocol, as a complement to the average success rate. On LIBERO-PRO, a public benchmark of perturbed LIBERO tasks for which no demonstrations exist, 27 of 57 scored tasks are at exactly 0% success for a pi_0.5 policy fine-tuned on the original LIBERO tasks. Direct PPO from this policy, under the same PPO recipe and the same RL compute as SynthDemo-RL's refinement stage, rescues 10 of these 27 tasks and leaves 17 at 0%. SynthDemo-RL, with 50 synthesized trajectories per task and no new human demonstrations, rescues all 27 and reaches average success rates of 97.8% and 97.1% on the Position and Task axes of LIBERO-PRO, respectively. On standard LIBERO, the same pipeline reaches 96.0% with no human demonstrations, within 1.7 points of pi_0.5 trained on 50 human demonstrations per task. We further validate the pipeline on RoboTwin 2.0 and verify that trajectories from a policy trained in a MuJoCo twin execute open-loop on a physical robot.

补充信息

↑