发表机构
Fudan University; Shanghai AI Laboratory; Harbin Institute of Technology, Shenzhen, China; Northwestern Polytechnical University; TeleAI, China Telecom Corp Ltd(复旦大学; 上海人工智能实验室; 哈尔滨工业大学(深圳); 西北工业大学; 中国电信股份有限公司TeleAI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MixVLA提出自适应混合非不变信息的训练框架,无需额外数据或架构改动,通过正则化分布特定变异性并融合不变特征,提升VLA模型在分布外场景的零样本泛化鲁棒性,同时保持域内性能。
AI 中文摘要
视觉-语言-动作(VLA)模型在机器人操作领域取得了显著进展,但其在分布外(OOD)条件下的零样本泛化能力仍然有限。这些模型常常将任务相关的不变结构与环境特定的非不变因素纠缠在一起,导致策略在动作预测时依赖虚假的外观线索。在本工作中,我们提出了MixVLA,一种模型无关的训练框架,无需额外的OOD数据或架构修改即可提升VLA模型的泛化能力。MixVLA的关键组件是自适应非不变信息混合(AMI)。AMI随机混合非不变表示,以正则化分布特定的变异性,同时保留互补的预测线索。混合后的非不变特征随后与不变表示融合,用于最终的动作预测,从而在不牺牲策略表达能力的情况下提高鲁棒性。在具有挑战性的操作场景中进行的广泛实验,包括LIBERO、LIBERO-Plus、RoboTwin扰动套件以及真实世界任务,表明MixVLA在保持强域内性能的同时,提升了整体的零样本鲁棒性。
英文摘要
Vision-Language-Action (VLA) models have achieved remarkable advances in robotic manipulation, yet their zero-shot generalization under out-of-distribution (OOD) conditions remains limited. These models often entangle task-relevant invariant structure with environment-specific non-invariant factors, causing policies to rely on spurious appearance cues during action prediction. In this work, we propose \textbf{MixVLA}, a model-agnostic training framework that improves the generalization of VLA models without requiring additional OOD data or architectural modifications. The key component of MixVLA is \textbf{Adaptive Mixing of Non-Invariant Information (AMI)}. AMI stochastically mixes non-invariant representations to regularize distribution-specific variability while preserving complementary predictive cues. The mixed non-invariant features are then fused with invariant representations for final action prediction, resulting in improved robustness without sacrificing policy expressiveness. Extensive experiments across challenging manipulation settings, including LIBERO, LIBERO-Plus, the RoboTwin perturbation suite, and real-world tasks, demonstrate that MixVLA improves overall zero-shot robustness while retaining strong in-domain performance.