arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PermVLA:分解顺序作为VLA学习的正则化器

PermVLA: Factorization Order as a Regularizer for VLA Learning

Yanqiao Chen, Yuhan Rui, Dongsheng Hou, Zijie Nie, Yutong Wan, Qi Hao

arXiv 2610.04659首次发表:更新:

发表机构

Southern University of Science and Technology; Huazhong University of Science and Technology(南方科技大学; 华中科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出因果锚定排列(CAP)作为VLA学习的正则化方法,通过采样动作揭示顺序并训练共享策略,在LIBERO、LIBERO-Plus和CALVIN上优于标准LTR训练,并扩展到扩散动作生成器。

AI 中文摘要

视觉-语言-动作(VLA)策略通常通过固定的从左到右(LTR)分解来学习动作块,尽管相同的专家轨迹分布允许许多有效的链式法则分解。我们将分解顺序识别为一个被忽视的正则化选择,并引入了因果锚定排列(CAP),它通过可调的按时间顺序前缀来采样动作揭示顺序。其辅助目标训练一个共享策略,从同一专家块的不同已知子集预测动作,而部署时保留确定性的LTR控制。我们称之为条件集增强:它从一个专家块创建多个条件预测问题,而不增加演示。这阻止了对普通教师强制所使用的单一按时间顺序前缀的依赖。受控实验表明,CAP在LIBERO和LIBERO-Plus上始终优于标准LTR训练,同样的优势出现在跨数据集的CALVIN评估中。一种衡量在两种揭示顺序下块联合对数似然之间的期望平方差的诊断方法验证了CAP训练内化了跨揭示顺序的一致性。这些发现将采样子集条件辅助目标定位为构建VLA正则化器的通用配方,并通过扩展到扩散动作生成器加以说明。

英文摘要

Vision-language-action (VLA) policies commonly learn action chunks through a fixed left-to-right (LTR) factorization, although the same expert trajectory distribution admits many valid chain-rule factorizations. We identify factorization order as an overlooked regularization choice and introduce causally anchored permutation (CAP), which samples action reveal orders with a tunable chronological prefix. Its auxiliary objective trains one shared policy to predict actions from different known subsets of the same expert chunk, while deployment retains deterministic LTR control. We call this conditional-set augmentation: it creates multiple conditional prediction problems from one expert chunk without adding demonstrations. This discourages reliance on the single chronological prefix used by ordinary teacher forcing. Controlled experiments show that CAP consistently outperforms standard LTR training on LIBERO and LIBERO-Plus, with the same advantage appearing in cross-dataset CALVIN evaluation. A diagnostic that measures the expected squared difference between a chunk's joint log likelihood under two reveal orders verifies that CAP training internalizes agreement across reveal orders. These findings position sampled subset-conditioned auxiliary objectives as a general recipe for constructing VLA regularizers, illustrated by an extension to diffusion action generators.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑