arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35603cs.LGstat.ML

控制几何拉直用于基于采样的潜在规划

Control-Geometry Straightening for Sampling-Based Latent Planning

  • Texas A&M University(德克萨斯农工大学)

机构由 AI 辅助整理,请以论文原文为准。

Ziang Fu, Ning Ning

中文总结 AI 辅助

本文提出控制几何拉直(CGS)辅助损失,通过匹配动作与潜在差异的余弦相似度,在多个控制环境中提升采样型规划器的成功率,并给出理论保证。

中文摘要 AI 辅助

联合嵌入预测架构使得利用潜在世界模型进行规划成为可能,但仅凭准确的转移预测并不能确保规划目标易于优化。我们引入了控制几何拉直(CGS),这是一种单一的辅助损失,通过直接拉直控制几何来实现采样高效的规划,从而学习对规划友好的表示。CGS仅利用来自像素-动作对的局部转移,将动作之间的成对余弦相似度与相应潜在差异之间的成对余弦相似度进行匹配。该损失可应用于端到端学习或预训练表示的世界模型架构。在线性动力学下,我们的理论分析将此目标与时间拉直以及整个规划范围内更平衡的终端成本曲率联系起来,为MPPI提供有限预算保证,为CEM提供局部收缩结果,并为梯度下降提供收敛界限。在四个控制环境和多个规划器中,CGS以更少的采样候选和细化步骤改善了规划,与LeWorldModel(LeWM)及其时间拉直变体(LeWM+TS)相比,成功率分别提高了最多20和12.6个百分点,采样型规划器每次更新使用128个候选。探针实验、与DINO-WM架构的比较以及规划器侧的消融实验阐明了潜在运动组织、状态依赖性和动态上下文如何塑造规划行为。因此,拉直控制几何使得在有限的规划预算下更容易找到好的动作序列。

英文摘要

Joint-embedding predictive architectures enable planning with latent world models, but accurate transition prediction alone does not ensure that the planning objective is easy to optimize. We introduce Control-Geometry Straightening (CGS), a single auxiliary loss that learns planner-friendly representations by directly straightening control geometry for sampling-efficient planning. CGS matches pairwise cosine similarities among actions to those among corresponding latent differences only using local transitions from pixel-action pairs. The loss can be applied across world-model architectures using end-to-end learned or pretrained representations. Under linear-dynamics, our theoretical analysis connects this objective to temporal straightening and more balanced terminal-cost curvature across the full planning horizon, yielding finite-budget guarantees for MPPI, local contraction results for CEM, and convergence bounds for gradient descent. Across four control environments and multiple planners, CGS improves planning with fewer sampled candidates and refinement steps, achieving success-rate gains up to 20 and 12.6 percentage points over LeWorldModel (LeWM) and its temporal-straightening variant (LeWM+TS), respectively, with sampling-based planners using 128 candidates per update. Probes, comparisons with DINO-WM architecture, and planner-side ablations clarify how latent motion organization, state dependence, and dynamical context shape planning behavior. Straightening control geometry thus makes good action sequences easier to find under limited planning budgets.

↑