发表机构
Texas A&M University; Purdue University(德克萨斯A&M大学; 普渡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出RoboIRS,一种无需更新参数、利用成败轨迹训练分类器并引导内部表示的推理时方法,在模拟和真实任务中显著提升VLA和世界-动作模型的成功率。
AI 中文摘要
视觉-语言-动作(VLA)模型和世界-动作模型(WAMs)在遇到分布外任务变化时,尽管保留了部分任务能力,性能仍常常下降。为恢复此类能力,我们提出RoboIRS,一种推理时内部表示引导方法,利用成功和失败的轨迹训练线性分类器,选择与结果相关的干预位置,并推导任务特定的引导方向,而无需更新策略参数。在15个模拟任务中,使用冻结的π0.5策略,RoboIRS将平均成功率从44.4%提升至66.2%,优于其他推理时干预基线,且仅增加少量推理时间。我们进一步在真实机器人操作中验证RoboIRS,使用相同的π0.5策略,并展示其适用于世界-动作模型Cosmos Policy,平均成功率从35.4%提升至55.4%。这些结果表明,直接引导机器人策略的内部表示可以在推理时提升机器人策略的性能。项目网站见本HTTPS链接。
英文摘要
Vision-language-action (VLA) and world-action models (WAMs) often degrade under out-of-distribution task variations despite retaining partial task capability. To recover such capability, we propose RoboIRS, an inference-time internal representation steering method that uses successful and failed rollouts to train linear classifiers, select outcome-relevant intervention locations, and derive task-specific steering directions without updating policy parameters. On 15 simulation tasks with a frozen $π0.5$ policy, RoboIRS improves the average success rate from 44.4% to 66.2%, outperforming alternative inference-time intervention baselines while adding little inference time. We further validate RoboIRS on real-robot manipulation using the same $π0.5$ policy and demonstrate its applicability to a world-action model Cosmos Policy, where the average success rate improves from 35.4% to 55.4%. These results show that directly steering internal robot-policy representations can improve the performance of robot policies at inference time. Project website is available at https://rollingoat.github.io/roboirs/.