发表机构
Graduate School of Engineering, The University of Tokyo(东京大学工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对SIGReg在多任务LeWM中导致表示混叠的问题,提出将其应用于时间中心残差的改进方法,在LIBERO基准上显著提升了多任务世界模型的下游性能。
AI 中文摘要
近期关于LeWorldModel(LeWM)的研究表明,草图各向同性高斯正则化器(SIGReg)通过将潜在边际分布正则化为各向同性高斯,能够实现从像素进行稳定的端到端世界模型学习,从而防止表示崩溃。尽管该方法在单任务设置中有效且简洁,但该方案无法可靠地扩展到多任务训练,导致下游行为克隆性能显著下降。本文中,我们发现边际高斯化会压缩任务相关潜在簇之间的分离度,相对于簇内变异而言。这种压缩会在任务和状态间引入表示混叠,并使学习到的表示对微小视觉扰动高度敏感。为解决该问题,我们将SIGReg应用于时间中心残差而非潜在边际分布。该替代目标不会对簇中心之间的分离施加直接正则化压力,消除了整个潜在空间需遵循单一各向同性高斯的要求,同时保留了SIGReg的抗崩溃效果。在LIBERO基准测试中,我们的方法提升了长 horizon 套件的下游成功率1.7倍,并将四个套件的平均成功率从53.2%提高到73.6%。在无外部预训练的情况下,其性能略优于从头训练的Diffusion Policy,并接近大规模预训练策略基线的性能。这些结果揭示了边际高斯先验与多任务潜在结构之间存在结构不兼容性,并为稳定且可扩展的端到端多任务世界模型学习提供了一条简单路径。
英文摘要
Recent work on LeWorldModel (LeWM) has shown that the Sketched Isotropic Gaussian Regularizer (SIGReg) enables stable end-to-end world model learning from pixels by regularizing the latent representation toward an isotropic Gaussian. While effective for latent-space planning, the representations learned by Raw LeWM are poorly suited for downstream robot policy learning. In this paper, through Monte Carlo analysis, we show that the Raw LeWM objective biases variance allocation toward the temporally persistent component, thereby suppressing the variance of the temporally centered residual. Consistent with this analysis, trained Raw LeWM representations exhibit suppressed residual variation and reduced decodability of robot state and dynamics, particularly gripper dynamics, which are crucial for robotic manipulation. To address this issue, we apply SIGReg to temporally centered residuals rather than to the whole latent representation. This simple change decouples persistent and residual variance allocation while retaining an effective anti-collapse property. On the LIBERO benchmark, our method improves downstream policy success on the Goal suite by 1.66x and raises the average success rate across all suites from 63.6% to 83.8%. Without external pretraining, it also outperforms both Diffusion Policy trained from scratch and the pretrained OpenVLA baseline. These results associate the variance-allocation bias of Raw LeWM with the downstream policy gap, and show that decoupling persistent and residual variation yields representations better suited for downstream robot policy learning.