用于减轻机器人基础模型中捷径学习的人工中央凹感知
Artificial Foveated Perception for Mitigating Shortcut Learning in Robotic Foundation Models
浏览论文内容
中文总结 AI 辅助
研究机器人基础模型因捷径学习难以稳健部署的问题,提出人工中央凹感知模块AFP,通过预测任务条件掩码辅助微调,使策略关注任务相关区域,实验证明其能减少微调时间、抑制过拟合并提升泛化能力。
中文摘要 AI 辅助
机器人基础模型在多任务能力等方面取得进展,但在不同现实场景中稳健部署仍困难,部分原因是策略常无法区分因果相关视觉结构与虚假场景相关性,即捷径学习。本文提出人工中央凹感知(AFP),它是轻量级、与策略无关的模块,以视觉和语言输入预测任务条件掩码。在微调时用作辅助基础信号,微调后策略在原始观测流上执行。实验表明AFP减少微调时间、抑制过拟合并改善泛化能力,消融实验显示这些收益源于引导策略学习到与任务相关的视觉证据。
英文摘要
Robotic foundation models still need task-specific fine-tuning before deployment, and the fine-tuned policies often break under modest changes in scene layout, lighting, or nearby distractors. We trace this brittleness to \textit{shortcut learning}: fine-tuning supervises actions but not the visual evidence the policy uses, so the policy can settle on scene-level correlations that predict the demonstrations without causing success. We propose Artificial Foveated Perception (AFP), a lightweight, policy-agnostic module that takes the same vision and language inputs as existing Vision-Language-Action and World Action Model pipelines and predicts task-conditioned masks over the relevant objects, the robot, and other action-critical regions. During fine-tuning the masks serve as an auxiliary grounding signal that aligns the policy's visual attention with task-relevant regions; the policy architecture is unchanged, and at inference the policy runs on the original observation stream with no AFP call in the control loop. In simulation with four robotic foundation models and on a real robot with $π_{0.5}$, AFP improves generalization under environmental perturbations, reduces overfitting, and shortens fine-tuning. Ablations over mask quality and grounding-loss design show that these gains come from directing policy learning toward task-relevant visual evidence. Code, data, and videos are available at https://apollo-lab-yale.github.io/26-CoRL-AFP-website/.
发表机构
- Yale University(耶鲁大学)
- University of Connecticut(康涅狄格大学)
- Peking University(北京大学)
- Imperial College London(帝国理工学院)
- Digients
机构由 AI 辅助整理,请以论文原文为准。