发表机构
University of Toronto(多伦多大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
RouteRLT学习何时及由哪个RL专家控制通用VLA策略,通过阶段选择器、稳定器和动作边界管理器实现精度关键任务的自动路由,在仿真和真实机器人上验证了其有效性和恢复能力。
AI 中文摘要
视觉-语言-动作(VLA)模型提供了广泛的操控能力,但在精度关键阶段常常表现不佳,而这些阶段主导了诸如连接器插接和线缆管理等高接触性工业任务。一种常见的补救措施是使用强化学习(RL)对预训练的VLA进行微调,从而在行为克隆之外实现任务特定的改进。然而,如何在决定何时需要RL微调以及哪个专门策略应执行的同时,保持其通用行为,仍是一个开放问题。在这项工作中,我们提出了RouteRLT,一个路由框架,它学习何时以及哪个RL专家(即为单一精度关键阶段训练的RL策略)应从通用VLA手中接管控制。一个阶段选择器识别当前控制器,一个稳定器抑制瞬态切换,一个动作边界管理器处理分块策略输出之间的转换。我们在LIBERO中的多物体抓取放置任务上评估了RouteRLT,以及一个具有多个精度关键阶段的真实世界线缆拾取和端口插接任务。在仿真中,学习到的路由优于基础VLA,并在不访问特权阶段边界的情况下,与使用特权阶段边界的路由相匹配。真实机器人评估验证了在操作员对齐交接协议下,自动路由到拾取和插接两个专家。总体而言,这些结果表明,学习到的路由在最需要精确适应的位置应用RL专家控制,同时保持通用VLA行为,包括从失败执行尝试中恢复。
英文摘要
Vision-language-action (VLA) models often struggle in the precision-critical phases of multi-stage manipulation tasks. To mitigate this issue, VLA models can be used in conjunction with reinforcement learning (RL) specialists that are specifically trained to handle the precision-critical phases. However, the coordination between the base VLA model and the RL specialists, which dictates when a specialist should take over from the base VLA and vice-versa, remains an open research question. In this paper, we address this gap by introducing RouteRLT, a modular framework that coordinates a generalist VLA, used as the default controller, with designated precision-critical RL specialists. At a high level, our framework trains a phase-aware coordination mechanism that handles handoffs between the generalist and the specialists. We evaluate RouteRLT on the LIBERO and LIBERO-Plus benchmarks, as well as on a physical connector pickup and insertion task. Overall, we find that RouteRLT improves success on LIBERO, and retains net gains on LIBERO-Plus. On the physical task, RouteRLT completes 65.7% of trials, compared with 8.6% for the baseline. Altogether, these results demonstrate that learned coordination builds on generalist VLA capabilities to improve task completion in precision-critical manipulation.
Comments8 pages, 5 figures, 4 tables. Extended version of the paper accepted at the IROS 2026 International Workshop on Industrial Applications of Robot Learning (IARL), with expanded multi-backbone evaluation, robustness analysis, and autonomous real-robot experiments