发表机构
Department of Computer Science, Rutgers University(罗格斯大学计算机科学系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对长时机器人操作中VLA策略的不足,提出前瞻性残差强化学习,通过离线估计前瞻性值优化交接质量,在螺母拧紧装配任务中,该方法全任务成功率高,优于标准方法和基线,强调塑造子任务成功状态对提升长时性能的重要性。
AI 中文摘要
视觉-语言-动作(VLA)策略提供了强大的通用操作先验,但由于长时信用分配和子任务耦合,在公差要求严格、接触丰富的装配任务中常常失败。在基于冻结VLA基础策略的残差强化学习中展示了这种失败模式。提出前瞻性残差强化学习,通过用离线估计的前瞻性值增强每个子任务的稀疏成功奖励来优化交接质量。在基于扳手的螺母拧紧装配任务中,该方法实现了85.6%的全任务成功率,优于标准子任务残差强化学习和VLA基线,同时每个子任务的成功率不变。结果表明,提高长时性能需要塑造每个子任务产生的成功状态,而不仅仅是成功是否发生。
英文摘要
Vision-Language-Action (VLA) policies offer strong general-purpose manipulation priors, but often fail on tight-tolerance, contact-rich assembly due to long-horizon credit assignment and subtask coupling: a state that is geometrically successful for the current skill can be brittle for downstream skills. We show this failure mode in residual reinforcement learning (RL) over a frozen VLA base policy: constant sparse success rewards improve each subtask in isolation yet yield little or no gain when skills are chained, because terminal state quality is uncontrolled. We propose Foresight Residual RL, which optimizes handoff quality by augmenting each subtask's sparse success reward with an offline-estimated foresight value -- the probability of future subtask success conditioned on the terminal state of the current subtask. Concretely, we (i) train a visual foresight predictor from images of terminal states of the base policy, labeled using downstream rollout statistics, and (ii) train residual policies via backward foresight induction, using the predictor output as a reward multiplier. On a three-phase wrench-based nut-tightening assembly task in Isaac Gym (grasp, move-insert, rotate), our method achieves 85.6% full-task success, outperforming standard subtask residual RL (54.5%) and VLA baselines, while leaving per-subtask success unchanged. These results highlight that improving long-horizon performance requires shaping which successful states are produced at each sub-task, not only whether success occurs.
CommentsAccepted at IROS2026. Project website: https://jaysparrow.github.io/foresight-residual-rl