发表机构
University College London (UCL)(伦敦大学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过配对安慰剂和简单回退基线,评估冻结VLA策略的运行时修正效果,发现学习型修正仅在特定任务上带来显著净收益,且部分挽救源于基础策略自身变异性。
AI 中文摘要
在部署针对冻结的视觉-语言-动作(VLA)策略的运行时恢复之前,必须确认干预措施在超越普通运行间变异方面能提升成功率,并且其复杂性相对于简单动作具有附加价值。我们在四个RoboTwin任务上对冻结的$\pi_{0.5}$评估了这些问题。对于每个测试种子,我们将有修正和无修正的rollout配对,并包含同一种子的基础策略重跑作为安慰剂。种子簇区间和预设比较规则用于评估相对于随机结果变化的净收益。在3,888个配对episode中,完整流程在beat_allowbreak block_allowbreak hammer任务上将成功率提升了$+13.5$个百分点(95%区间$[+9.4,+17.7]$),而在部署权重下,其他三个任务没有检测到明显收益。在响应任务中失败的基础episode中,$43.2\\%$在简单重跑中成功,而修正后为$63.5\\%$;因此,许多名义上的挽救反映了基础策略自身的变异性。固定时间触发器和脚本化返回到较早的关节配置产生了净收益,并且在两轮中与学习型流程没有检测到差异,尽管我们的预设等价标准并未始终满足。暂停和恒定动作控制没有产生相当的收益。在此基准上,是否干预的决策强烈依赖于任务,并且需要配对安慰剂和简单回退基线来确定学习型修正的贡献。
英文摘要
Before deploying runtime recovery for a frozen vision-language-action (VLA) policy, one must establish that an intervention improves success beyond ordinary run-to-run variation and that its complexity adds value over a simple action. We evaluate these questions on frozen $π_{0.5}$ across four RoboTwin tasks. For each test seed, we pair rollouts with and without correction and include a same-seed base-policy re-run as a placebo. Seed-cluster intervals and prespecified comparison rules assess net gains against stochastic outcome changes. Across 3,888 paired episodes, the full pipeline raises success on beat_allowbreak block_allowbreak hammer by $+13.5$\,pp (95\% interval $[+9.4,+17.7]$), with no detectable gain on the other three tasks at the deployed weight. Among failed base episodes on the responsive task, $43.2\%$ succeed on a plain re-run, compared with $63.5\%$ after correction; many nominal rescues therefore reflect the base policy's own variability. A fixed-time trigger and scripted return to an earlier joint configuration produce a net gain with no detected difference from the learned pipeline across two rounds, although our prespecified equivalence criterion is not met consistently. Pausing and a constant-action control do not yield comparable gains. On this benchmark, the decision to intervene depends strongly on the task, and a paired placebo plus a simple retreat baseline are needed to establish what learned correction contributes.
Comments15 pages, 6 figures