arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GDLAM:用于高度解耦具身预训练的组解耦潜动作模型

GDLAM: Group-Disentangled Latent Action Model for Highly Disentangled Embodied Pretraining

Jiarui Yang, Jiawei Li, Jiale Zhang, Hang Guo, Wen Huang, Maowei Hu, Tao Dai, Shu-Tao Xia

arXiv 2609.26118首次发表:更新:

AI 中文总结

提出组解耦潜动作模型GDLAM,通过分组变分瓶颈和门控路由实现因果解耦,提升解耦指标且不损失动作信息,并在世界模型和VLA策略上取得更优性能。

AI 中文摘要

潜动作模型(LAMs)通过自监督未来预测从无动作视频中学习与动作相关的表示,为具身智能预训练提供了一种可扩展的范式。然而,现有的LAMs将视觉变化的异质来源(包括相机运动、物体动态和交互事件)压缩到单个潜向量中,导致表示纠缠且语义结构有限,从而限制了世界模型的可控性和VLA策略的泛化能力。我们提出了组解耦潜动作模型(GDLAM),这是一种潜动作模型,其编码通过构造被分解为N个组,每个组具有独立的变分瓶颈和空间门控路由路径,并通过一组信息几何目标进行训练:互斥性、组和门控稀疏性以及静态-动态正交性,这些目标使各组在因果上相互区分,而不仅仅是去相关。定量上,对任何单个组的干预仅改变该组,而保持其他组不变,并且GDLAM在无标签解耦指标(包括Modularity、MIG和DCI)上大幅优于强非结构化LAM。值得注意的是,这种分解不以动作信息为代价:在三个互信息估计器和一个线性探针上,分组编码在分布内和分布外都比整体基线更具信息量。作为解耦编码是可复用预训练货币的支持证据,我们进一步将其迁移到两个下游场景:(1)世界建模:使用GDLAM预训练的世界模型在 rollout 保真度和动作跟随能力上优于SOTA基线。(2)VLA策略:使用GDLAM预训练在多个仿真基准和真实世界机器人操作任务中显著提高了任务成功率,优于先前方法。

英文摘要

Latent action models (LAMs) learn action-related representations from action-free videos via self-supervised future prediction, offering a scalable paradigm for embodied intelligence pretraining. However, existing LAMs collapse heterogeneous sources of visual change, including camera motion, object dynamics, and interaction events, into a single latent vector, resulting in entangled representations with limited semantic structure and consequently restricting world model controllability and VLA policy generalization. We introduce the Group-Disentangled Latent Action Model (GDLAM), a latent action model whose code is factorized by construction into N groups, each with an independent variational bottleneck and a spatially gated routing pathway, and trained with a set of information-geometric objectives: mutual exclusivity, group and gate sparsity, and static-dynamic orthogonality, that make the groups mutually causally distinct rather than merely decorrelated. Quantitatively, intervening on any single group changes only that group and leaves the others intact, and GDLAM improves label-free disentanglement metrics, including Modularity, MIG, and DCI, by wide margins over a strong unstructured LAM. Notably, this factorization is not at the expense of action information: across three mutual-information estimators and a linear probe, the grouped code is more informative than monolithic baselines both in- and out-of-distribution. As supporting evidence that the disentangled code is a reusable pretraining currency, we further transfer it to two downstream regimes: (1) World Modeling: World models pretrained with GDLAM achieve superior rollout fidelity and action-following capability compared with SOTA baselines. (2) VLA Policies: Pretraining with GDLAM substantially improves task success rates over previous methods across multiple simulation benchmarks and real-world robotic manipulation tasks

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑