发表机构
Skylark Labs; Carnegie Mellon University; University of California, Berkeley(斯凯拉克实验室; 卡内基梅隆大学; 加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对部署后物理AI需应对新情况且不丢失原有能力的需求,提出CFAM模型,通过互补学习架构实现少样本技能学习与自主领域能力增长,在多实体评估中表现优于基线模型。
AI 中文摘要
无人交互式自主系统——即替代人类进入危险环境并使用人类工具完成任务的机器——仍是关键任务作业中缺失的能力。这些领域训练数据稀缺且仅能依赖板载计算资源,而部署后的系统必须在不丢失原有能力的前提下应对新情况。我们提出持续领域自适应模型(Continual Field-Adaptive Models, CFAMs),该模型可在实验室中高效学习,并通过自主、无梯度的设备端更新在部署后持续学习。CFAM采用互补学习架构,包含一个冻结的慢学习组件和一个快速学习胶囊域(Capsule Field)。慢学习组件包含三个皮层:传感器(Sensor),将多模态输入映射为基于三维空间的几何结构;推理(Reasoning),将任务分解为技能并评估结果;动作(Action),执行几何技能。胶囊域以能力胶囊(Competence Capsules)的形式一次性且无梯度地存储领域学习内容。技能安装在实验室中为少样本学习,在部署后为持续学习;开放世界新情况不在本研究范围内。我们在五种实体上评估CFAM:机械臂、四足机器人、人形机器人、四旋翼无人机和越野车辆。基线模型(pi0、CogACT、SpatialVLA)使用相同的内部多实体数据集进行物理平台对比。CFAM使用40%的数据(即少2.5倍的轨迹)即可达到在全部先前训练数据集上训练的标准策略的操作点。在测试时,自主捕获经验证的近分布外(near-OOD)案例使动作成功率提高13.9个百分点。在顺序仿真中,其反向迁移为-0.5个百分点,而LoRA为-11.4个百分点。因此,CFAM提供了一种有界形式的部署后物理智能:少样本技能学习、从经验证的近OOD经验中自主实现领域能力增长,以及保留原有能力。
英文摘要
Unattended interactive autonomy - machines that step into danger in place of humans and complete tasks with human tools - remains a missing capability in mission-critical operations. These domains offer scarce training data and only onboard compute, yet deployed systems must face novelty without erasing prior competence. We introduce Continual Field-Adaptive Models (CFAMs), which learn efficiently in the lab and continue learning after deployment through autonomous, gradient-free, on-device updates. CFAM uses a complementary learning architecture with a frozen slow-learning component and a fast-learning Capsule Field. The slow component contains three cortices: Sensor, which maps multimodal input into 3D-grounded geometry; Reasoning, which decomposes tasks into skills and evaluates outcomes; and Action, which executes geometric skills. The Capsule Field stores field learning one-shot and gradient-free as Competence Capsules. Skill installation is few-shot in the lab and continual in the field; open-world novelty is outside scope. We evaluate CFAM across five embodiments: manipulator, quadruped, humanoid, quadrotor, and off-road vehicle. Baselines (pi0, CogACT, SpatialVLA) use the same in-house multi-embodiment dataset for physical-platform comparisons. CFAM reaches the operating point of a standard policy trained on the full prior-training dataset using 40% of the data, or 2.5x fewer trajectories. At test time, autonomous capture of verified near-OOD cases improves action success by 13.9 percentage points. In sequential simulation, backward transfer is -0.5 percentage points versus -11.4 for LoRA. CFAM therefore provides a bounded form of post-deployment physical intelligence: few-shot skill learning, autonomous field growth from verified near-OOD experience, and retention of prior competence.