发表机构
South China University of Technology; National University of Singapore; Nanyang Technological University(华南理工大学; 新加坡国立大学; 南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对机器人基础模型在视觉分布偏移下因视觉-动作捷径而泛化失败的问题,提出两阶段潜在接口训练(LIT),通过姿态监督约束视觉条件化,在多种架构上显著提升LIBERO-Plus及真实世界任务成功率。
AI 中文摘要
机器人基础模型在分布内任务上表现优异,但在视觉分布偏移下性能往往下降。当从预训练的视觉表征中学习生成动作时,模型可能利用训练分布中与示范动作相关的任务无关视觉线索。当这些相关性在分布偏移下发生变化时,此类视觉-动作捷径会削弱泛化能力。缓解这些捷径需要在保留任务相关空间信息的同时,约束视觉信息用于动作生成的方式。我们提出潜在接口训练(Latent Interface Training, LIT),一种框架无关的两阶段策略:首先在无图像条件下建立空间目标条件化的动作先验,然后通过姿态监督的潜在接口约束视觉条件化。第一阶段训练动作专家,使其基于语言、机器人状态以及每个示范片段末端的SE(3)末端执行器姿态生成动作块,从而独立于视觉线索学习目标导向的动作生成。第二阶段引入一个潜在接口,聚合视觉和语义表征,并作为预训练动作专家唯一的视觉条件化通路。该接口受监督以重建第一阶段用于条件化的末端姿态,促使其保留动作生成所需的目标相关空间信息。在四种视觉-语言-动作及世界-动作架构(Pi0.5、MolmoAct2、FAST-WAM和ImageWAM)上,LIT将LIBERO-Plus整体成功率提升了3.87至10.70个百分点,同时保持或提升了LIBERO平均成功率。真实世界评估显示,在未见过的相机配置、光照变化和干扰物条件下,三个任务的成功率总计提升了13.30至16.70个百分点。
英文摘要
Robot foundation models achieve strong in-distribution performance but often degrade under visual distribution shifts. When learning to generate actions from pretrained visual representations, models may exploit task-irrelevant visual cues that correlate with demonstrated actions within the training distribution. Such vision-action shortcuts can undermine generalization when these correlations change under distribution shifts. Mitigating these shortcuts requires constraining how visual information is used for action generation while preserving task-relevant spatial information. We propose Latent Interface Training (LIT), a framework-agnostic two-stage strategy that first establishes a spatial-goal-conditioned action prior without images, then constrains visual conditioning through a pose-supervised latent interface. Stage 1 trains the action expert to generate action chunks conditioned on language, robot state, and each demonstrated chunk's terminal SE(3) end-effector pose, learning goal-directed action generation independently of visual cues. Stage 2 introduces a latent interface that aggregates visual and semantic representations and serves as the pretrained action expert's only visual conditioning pathway. The interface is supervised to reconstruct the terminal pose previously used to condition Stage 1, encouraging it to retain the goal-relevant spatial information needed for action generation. Across four vision-language-action and world-action architectures (Pi0.5, MolmoAct2, FAST-WAM, and ImageWAM), LIT improves overall LIBERO-Plus success by 3.87-10.70 percentage points while preserving or improving average LIBERO success. Real-world evaluations show 13.30-16.70 percentage-point gains in success aggregated across three tasks under unseen camera configurations, lighting variations, and distractors.
CommentsProject page: https://magiclab-nus.github.io/LIT/?v=37955c5