FAN:面向视觉-语言-动作模型持续适应的预见性动作归一化
FAN: Foresight Action Normalization for Continual Adaptation of Vision-Language-Action Models
另 2 家 · 查看机构详情
- HKU(香港大学)
- SUSTech(南方科技大学)
- EIT, Ningbo(宁波东方理工大学)
- HUST(华中科技大学)
- INFIFORCE
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对VLA模型持续适应中动作归一化被忽视的问题,提出预见性动作归一化(FAN),基于一致性、覆盖性和因果性三原则,在持续学习前从任务无关校准集估计并冻结归一化统计量,在多个真实任务流中取得最高性能。
中文摘要 AI 辅助
在大规模封闭数据集上预训练的视觉-语言-动作(VLA)模型已在多种机器人操作任务中展现出显著成功。然而,其长期现实世界部署需要不断获取新技能,同时保留先前学到的能力。尽管开创性工作已利用经验回放和强化微调等技术探索了持续的VLA适应,但它们忽略了一个基础机制:动作归一化,该机制决定了策略感知和执行物理动作所依据的底层坐标系。为弥补这一空白,我们系统评估了五种归一化策略,涵盖四个真实世界任务流,包括单臂和双臂操作。我们的分析揭示,现有协议因任务间坐标漂移、运动覆盖有限或训练-测试坐标不匹配而引发严重失效模式。受这些见解启发,我们制定了三个核心设计原则:一致性、覆盖性和因果性(3C),并引入了预见性动作归一化(FAN)。FAN在持续学习之前,从一个小型、任务无关的校准集中估计归一化统计量,并在整个适应过程中将其冻结。在所有评估的任务流中,FAN实现了最高性能并展现出持续稳健性,为构建稳定的动作表示以实现有效的终身VLA适应提供了富有洞察力的指导。
英文摘要
Vision-Language-Action (VLA) models pre-trained on large-scale, closed datasets have demonstrated remarkable success across diverse robotic manipulation tasks. However, their long-term real-world deployment necessitates continuously acquiring new skills while retaining previously learned capabilities. While pioneering works have explored continual VLA adaptation using techniques such as experience replay and reinforcement fine-tuning, they overlook a foundational mechanism: action normalization, which determines the underlying coordinate system in which policies perceive and execute physical actions. To bridge this gap, we systematically evaluate five normalization strategies across four real-world task streams covering single-arm and bimanual manipulation. Our analysis reveals that existing protocols induce severe failure modes due to inter-task coordinate drift, limited motion coverage, or train-test coordinate mismatches. Motivated by these insights, we formulate three core design principles: consistency, coverage, and causality (3C), and introduce foresight action normalization (FAN). FAN estimates normalization statistics once from a small, task-independent calibration set prior to continual learning and freezes them throughout adaptation. Across all evaluated streams, FAN achieves the highest performance and demonstrates consistent robustness, providing insightful guidance for building stable action representations in achieving effective lifelong VLA adaptation.