arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33595cs.ROcs.LG

超越单步精度:状态仿射潜在转移用于可靠的视觉规划

Beyond One-Step Accuracy: State-Affine Latent Transition for Reliable Visual Planning

Boyuan Zhang, Yingjun Du, Xiantong Zhen, Ling Shao

首次发表
浏览论文内容

中文总结 AI 辅助

针对视觉规划中单步训练导致递归误差传播的问题,提出状态仿射潜在转移模型SALT,通过递归多步监督训练,虽单步误差增大,但显著提升闭环成功率并降低失败风险。

中文摘要 AI 辅助

联合嵌入世界模型通过在潜在空间中学习动作条件下的动力学来实现视觉规划。然而,它们通常针对编码状态的单步预测进行训练,而规划则递归地将学习到的转移应用于其自身的预测。因此,单步精度并不能反映预测误差在递归展开下如何传播。我们将多步展开误差分解为各步骤引入的误差及其通过后续转移的传播。我们表明,状态仿射动力学正是具有状态无关雅可比矩阵的可微转移,消除了非线性传播残差,使得误差传播算子仅依赖于动作序列。受此结果启发,我们引入了SALT(状态仿射潜在转移),一种动作条件下的状态仿射动力学模型,其中动作同时调制状态变换和加性更新。我们通过递归多步展开监督来训练SALT,将每个预测的潜在状态反馈到转移中,使得训练与规划期间模型的使用方式相匹配。在四个视觉规划环境中,SALT的单步预测误差比匹配的LeWM基线高1.48至2.19倍,但在每个环境中的闭环成功率平均提高了10.0个百分点。在OGBench-Cube上,执行后模型预测成本急剧上升的失败片段比例从23.3%降至2.0%。

英文摘要

Joint-embedding world models enable visual planning by learning action-conditioned dynamics in latent space. Yet they are commonly trained for one-step prediction on encoded states, while planning recursively applies the learned transition to its own predictions. One-step accuracy therefore does not capture how prediction errors propagate under recursive rollout. We decompose multi-step rollout error into the errors introduced at individual steps and their propagation through subsequent transitions. We show that state-affine dynamics are precisely the differentiable transitions with state-independent Jacobians, eliminating the nonlinear propagation residual and making the error propagation operators depend only on the action sequence. Guided by this result, we introduce SALT (State-Affine Latent Transition), an action-conditioned state-affine dynamics model in which the action modulates both the state transformation and the additive update. We train SALT through recursive multi-step rollout supervision, feeding each predicted latent state back into the transition so that training matches how the model is used during planning. Across four visual planning environments, SALT exhibits $1.48$--$2.19\times$ higher one-step prediction error than the matched LeWM baseline, yet improves closed-loop success in every environment by $10.0$ percentage points on average. On OGBench-Cube, the fraction of episodes that fail with a sharp rise in model-predicted cost after execution decreases from $23.3%$ to $2.0%$.

发表机构

  • University of Chinese Academy of Sciences(中国科学院大学)
  • University of Amsterdam(阿姆斯特丹大学)
  • United Imaging Healthcare, Co., Ltd.(联影医疗科技股份有限公司)

机构由 AI 辅助整理,请以论文原文为准。

↑