发表机构
University of Southern California(南加州大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SWAP提出了一种离线强化学习驱动的动态策略路由框架,使VLA机器人能在执行中按步骤选择最优策略,在真实和模拟任务中成功率提升最高33%,动作步长减少28.3%。
AI 中文摘要
使用视觉-语言-动作(VLA)模型骨干的机器人操作系统通常仅使用一个VLA来执行任务。然而,单个VLA在不同任务状态和环境中的表现并不理想。我们提出了一种在执行过程中动态组合多个VLA策略的框架:逐步动作策略路由(SWAP)。SWAP将策略路由形式化为一个离线强化学习问题,学习一个路由评论员,该评论员根据当前观测在每个决策步骤选择最合适的策略。SWAP使机器人能够在在线执行过程中选择新的策略,而不是在整个回合中固定使用单一策略。我们在真实世界的DROID操作任务和LIBERO模拟实验上评估了SWAP。SWAP优于固定策略执行和路由基线,在真实世界任务成功率上实现了最高33%的绝对提升,同时将成功轨迹的机器人动作步长减少了28.3%。
英文摘要
Robot manipulation systems using Vision-Language-Action (VLA) model backbones typically use just one VLA for task execution. However, individual VLAs do not perform well across different task states and environments. We introduce a framework for dynamically composing multiple VLA policies during execution: StepWise Action Policy Routing (SWAP). SWAP formulates policy routing as an offline reinforcement learning problem, learning a routing critic that selects the most appropriate policy at each decision step given the current observation. SWAP enables robots to select new policies to execute online rather than committing to a single policy for the duration of an episode. We evaluate SWAP on both real-world DROID manipulation tasks and LIBERO simulation experiments. SWAP improves over fixed-policy execution and routing baselines, giving absolute improvements in real-world task success up to 33% while reducing successful trajectory robot action step length by 28.3%.
Comments9 pages , 4 figures,Under review for ICRA 2027