GIFT:面向机器人操作的基于动作导向结构监督的引导式中间特征训练
GIFT: Guided Intermediate Feature Training via Action-Oriented Structural Supervision for Robotic Manipulation
浏览论文内容
中文总结 AI 辅助
该研究提出 GIFT 框架,通过几何对齐、可供性预测和目标区域重建约束中间特征,在 LIBERO-Plus 和 RoboCasa 数据集上,其多个变体在机器人操作任务中显著优于基准模型,性能提升明显。
中文摘要 AI 辅助
视觉-语言预训练和预测世界模型为机器人策略提供了丰富的语义和动态视觉特征,但它们原生的动作和视觉预测目标可能会忽略关键的物理和任务结构,同时保留与控制无关的视觉冗余。我们将这种视觉丰富性与控制效用之间的不匹配称为动作充分性差距。我们研究是否可以通过引导中间特征保留机器人操作中三种与控制相关的结构来弥合这一差距:决定运动可行性的几何结构、编码指令相关实体的 affordance(可供性)、以及在任务相关区域中基于指令的目标。为此,我们提出了 GIFT(Guided Intermediate Feature Training,引导式中间特征训练),这是一种架构灵活的框架,用于学习中间特征,通过几何对齐、可供性预测和目标区域重建将这些结构转化为训练时的约束条件。我们在视觉-语言-动作(VLA)策略、直接动作世界模型(WAM)以及逆动力学 WAM 中实例化了 GIFT,同时保留了每个模型的动作形式。在零样本迁移至 LIBERO-Plus 时,GIFT-VLA、GIFT-WAM-Fast 和 GIFT-WAM-IDM 的性能分别优于 StarVLA-OFT、Fast-WAM 和 Fast-WAM-IDM 4.6、12.6 和 5.2 个百分点,分别达到 79.6%、72.6% 和 87.8%。在 RoboCasa 上,三种 GIFT 变体分别达到 61.4%、83.6% 和 82.3%,比对应的基准模型分别高出 12.6、9.0 和 8.4 个百分点。这些结果共同表明,学习具有功能结构的中间特征是跨模型特定动作形式的可复用原则,在关节物体任务和存在未见过的视觉与空间扰动的高精度现实世界操作中,性能提升尤为显著。项目页面:this https URL
英文摘要
Vision-language pre-training and predictive world modeling provide robot policies with rich semantic and dynamic visual features, but their native action and visual-prediction objectives may omit critical physical and task structure while retaining control-irrelevant visual redundancy. We call this mismatch between visual richness and control utility the action-sufficiency gap. We investigate whether this gap can be bridged by guiding intermediate features to preserve three control-relevant structure in robotic manipulation: geometry governing motion feasibility, affordance encoding instruction-relevant entities, and goals grounding instructions in task-relevant regions. To this end, we present GIFT (Guided Intermediate Feature Training), an architecture-flexible framework for learning intermediate features that translates these structures into training-time constraints through geometry alignment, affordance prediction, and goal-region reconstruction. We instantiate GIFT in a Vision-Language-Action (VLA) policy, a direct-action World-Action Model (WAM), and an inverse-dynamics WAM while retaining each model's action formulation. Under zero-shot transfer to LIBERO-Plus, GIFT-VLA, GIFT-WAM-Fast, and GIFT-WAM-IDM outperform StarVLA-OFT, Fast-WAM, and Fast-WAM-IDM by 4.6, 12.6, and 5.2 points, reaching 79.6%, 72.6%, and 87.8%, respectively. On RoboCasa, the three GIFT variants reach 61.4%, 83.6%, and 82.3%, outperforming their counterparts by 12.6, 9.0, and 8.4 points, respectively. Together, these results establish learning functionally structured intermediate features as a reusable principle across model-specific action formulations, with especially large gains on articulated-object tasks and high-precision real-world manipulation under unseen visual and spatial perturbations. Project page: https://openphoenix-team.github.io/GIFT-pages.
发表机构
- Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
- University of Chinese Academy of Sciences(中国科学院大学)
- National University of Singapore(新加坡国立大学)
- Tsinghua University(清华大学)
- Fudan University(复旦大学)
机构由 AI 辅助整理,请以论文原文为准。