发表机构
UNIST(蔚山科学技术院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出SALT方法,利用未执行动作块作为自监督信号,在测试时适应未知视觉干扰,显著提升VLA策略在模拟和真实机器人上的成功率与任务进度。
AI 中文摘要
在机器人执行任务过程中,视觉干扰可能随时出现,使得视觉-语言-动作(VLA)策略在不知道干扰类型或发生时机的情况下做出响应。我们提出了基于剩余轨迹的自监督适应方法(SALT),该方法利用剩余轨迹——即先前动作块中未执行的部分——作为测试时自监督信号。由于连续的动作块在时间上存在重叠,剩余轨迹为当前预测在相同未来控制区间内提供了时间对齐的目标。在视觉偏移发生时,剩余轨迹可能保留在干扰发生前形成的计划,因此将策略更新向该目标对齐,可锚定跨偏移的适应过程(过渡锚定)。SALT保留适应后的策略并重新生成当前动作块,其剩余轨迹在下次重新规划时成为新的目标,从而将修正沿执行轨迹向前传递(顺序修正传播)。监督信号完全来自策略自身的预测,无需干扰标注、专家动作或目标域演示,且仅基于名义轨迹校准的轻量级适应门控决定何时开始更新。在LIBERO-10基准上,SALT将五种持久视觉干扰下的平均成功率从43.9%提升至53.2%(使用SmolVLA),并从58.7%提升至66.0%(使用GR00T N1.7),同时基本保持名义性能。在真实机器人上,它将数字和物理干扰下的平均任务进度从0.49提升至0.61。
英文摘要
Visual disruptions can arise while a robot is executing a task, leaving a vision-language-action (VLA) policy to respond without knowing the disruption type or timing. We introduce Self-supervised Adaptation from Leftover Trajectories (SALT), which uses the leftover trajectory, the unexecuted part of the previous action chunk, as self-supervision for test-time adaptation. Because consecutive chunks overlap in time, the leftover provides a temporally aligned target for the current prediction over the same future control interval. At the onset of a visual shift, the leftover can retain a plan formed before the corruption, so updating the policy toward it anchors the adaptation across the shift (Transition Anchoring). SALT keeps the adapted policy and regenerates the current chunk, whose leftover becomes the target at the next replan, carrying the correction forward along the execution trajectory (Sequential Correction Propagation). Supervision comes entirely from the policy's own predictions, requiring no disruption annotations, expert actions, or target-domain demonstrations, and a lightweight adaptation gate calibrated only on nominal trajectories decides when updates begin. On LIBERO-10, SALT increases average success across five persistent visual corruptions from 43.9% to 53.2% with SmolVLA and from 58.7% to 66.0% with GR00T N1.7, while largely preserving nominal performance. On a real robot, it raises task progress averaged over digital and physical disruptions from 0.49 to 0.61.
Comments22 pages, 15 figures, 13 tables