arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

库普曼梦想家:用于稳定世界模型想象的频谱约束潜在动力学

Koopman Dreamer: Spectrally Constrained Latent Dynamics for Stable World-Model Imagination

Jiaqi Li, Xinglong Zhang, Haibin Xie, Yixing Lan, Wei Pan, Xin Xu

arXiv 2607.19719首次发表:更新:

发表机构

College of Intelligence Science and Technology, National University of Defense Technology; School of Engineering, Newcastle University(国防科技大学智能科学与技术学院; 纽卡斯尔大学工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对潜在世界模型长期展开中模态持久性和误差积累控制有限的问题,提出库普曼梦想家模型,通过频谱约束潜在动力学核心及多种目标结合优化,推导误差界,实验证明其提高了长期潜在展开稳定性及闭环控制性能。

AI 中文摘要

潜在世界模型通过在想象的潜在轨迹上优化策略来提高连续控制中的样本效率,但常见的神经转换在长期展开中对模态持久性和误差积累的直接控制有限。我们提出了库普曼梦想家,这是一种具有频谱约束确定性潜在动力学核心的梦想家式世界模型。其受库普曼启发的主干使用具有有界半径的二维旋转缩放块来表示阻尼、旋转和近周期模式。线性和低秩双线性作用项捕捉全局和状态依赖的控制效果,而随机状态调制提供局部校正信息。为减少后验条件训练与仅先验想象之间的不匹配,该模型将后验条件指数移动平均教师目标与单步一致性、多步展开和开环观测预测目标相结合。我们进一步推导了一个多步展开误差界,将频谱主干和双线性相互作用的放大与随机状态不匹配和建模残差的加性效应分开,阐明了误差衰减与长期信息保留之间的权衡。在深度思维控制套件的本体感觉连续控制任务和无人机激光雷达自主导航上的实验结果表明,库普曼梦想家提高了长期潜在展开的稳定性,并在依赖高质量多步想象的任务上实现了更强的闭环控制性能。

英文摘要

Latent world models improve sample efficiency in continuous control by optimizing policies over imagined latent trajectories, but common neural transitions offer limited direct control over modal persistence and error accumulation in long rollouts. We propose Koopman Dreamer, a Dreamer-style world model with a spectrally constrained deterministic latent dynamics core. Its Koopman-inspired backbone uses two-dimensional rotation--scaling blocks with bounded radii to represent damping, rotation, and near-periodic modes. Linear and low-rank bilinear action terms capture global and state-dependent control effects, while stochastic-state modulation supplies local correction information. To reduce the mismatch between posterior-conditioned training and prior-only imagination, the model combines posterior-conditioned EMA teacher targets with one-step consistency, multi-step rollout, and open-loop observation-prediction objectives. We further derive a multi-step rollout-error bound that separates amplification by the spectral backbone and bilinear interaction from the additive effects of stochastic-state mismatch and modeling residuals, clarifying the trade-off between error attenuation and long-term information retention. Experimental results on proprioceptive continuous-control tasks from the DeepMind Control Suite and UAV-LiDAR autonomous navigation demonstrate that Koopman Dreamer improves the stability of long-horizon latent rollouts and achieves stronger closed-loop control performance on tasks that rely on high-quality multi-step imagination.

Comments20 pages, 13 figures, 11 tables. Revised manuscript with a more concise and precise abstract and improved clarity and presentation throughout the main text. The main technical content, experimental results, and conclusions remain unchanged

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑