发表机构
Apple(苹果公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
REFACTOR-VLA 是用于学习可复用技能的醒/睡系统,在 LIBERO 数据集上,其改进的训练目标提升了聚类性能,生成了首个 real-LIBERO 任务-语言库,性能优于现有基线。
AI 中文摘要
大多数视觉-语言-动作(VLA)模型——OpenVLA、π₀、RT-2、RDT-1B——都是整体式的:它们输出原始运动指令或短动作块,却未将行为组织成可复用的抽象概念,因此在长 horizon 任务上性能下降且难以解释。现有的技能发现方法回避了“两个动作序列何时在行为上等价”这一核心问题,要么对对比嵌入进行聚类,要么将判断权交给未针对机器人动力学校准的语言模型。我们提出 REFACTOR-VLA,一种用于学习可复用技能的醒/睡系统。其“睡阶段”基于从学习到的潜在世界模型 M_φ 的 rollouts 计算出的行为等价核(BEK)对运动程序片段进行聚类;其“醒阶段”输出基于 Hindley–Milner 式词汇表的类型化 lambda 项,由库条件整流流动作解码器使用。抽象概念仅在通过最小描述长度和返回保留门时才被接纳。在 LIBERO 上我们报告两项发现:第一,将世界模型参数从 1.88 亿扩大到 4.3 亿,会使 4 个套件中 4 个的性能下降,因此仅靠容量无法提升性能;第二,训练目标的重要性远高于此:在世界模型预热期间添加辅助监督对比(InfoNCE)损失,可显著改善睡阶段聚类,在 n=3 个随机种子下,归一化互信息为:对象类 0.462±0.021、空间类 0.867±0.025、目标类 0.915±0.013、LIBERO-10 类 0.754±0.010,且在所有 4 个套件上比最强的已发表基线平均高出 Δ=+0.184。在 12 个提供方中,平均成对 NMI 的 95%自助法置信区间为 [0.683, 0.729](均值 0.705)。睡阶段还生成了首个 real-LIBERO 任务-语言库:解码器使用了 3 个已接纳抽象概念中的 2 个,并改写了所有 256 个采样演示。
英文摘要
Most vision-language-action (VLA) models -- OpenVLA, $π_0$, RT-2, RDT-1B -- are monolithic: they emit raw motor commands or short action chunks without organizing behavior into reusable abstractions, so they degrade on long-horizon tasks and resist interpretation. Existing skill-discovery methods sidestep the core question of when two action sequences are behaviorally equivalent, either clustering contrastive embeddings or delegating the judgment to a language model uncalibrated to the robot's dynamics. We introduce REFACTOR-VLA, a wake/sleep system for learning reusable skills. Its sleep phase clusters motor-program fragments under a Behavioral-Equivalence Kernel (BEK) computed from rollouts of a learned latent world model $M_ϕ$; its wake phase emits typed lambda terms over a Hindley--Milner-inspired vocabulary, consumed by a library-conditioned rectified-flow action decoder. Abstractions are admitted only if they pass Minimum Description Length and return-preservation gates. On LIBERO we report two findings. First, enlarging the world model from 188M to 430M parameters worsened performance on 4 of 4 suites, so capacity alone does not help. Second, the training objective matters far more: adding an auxiliary supervised contrastive (InfoNCE) loss during world-model warmup substantially improves sleep-phase clustering, giving Normalized Mutual Information at $n=3$ seeds of $0.462 \pm 0.021$ (object), $0.867 \pm 0.025$ (spatial), $0.915 \pm 0.013$ (goal) and $0.754 \pm 0.010$ (LIBERO-10), and beating the strongest published baseline on all 4 suites by a mean $Δ= +0.184$. Across providers ($n=12$) the 95% bootstrap confidence interval for mean pairwise NMI is $[0.683, 0.729]$ (mean $0.705$). The sleep phase also yields the first real-LIBERO task-language library: the decoder uses 2 of 3 admitted abstractions and rewrites all 256 sampled demonstrations.
Comments30 pages, 5 figures