arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.07265math.OCcs.LG

用于控制的学习型世界模型中的度量非崩溃:逼近理论、有限样本几何保证与确定性规划迁移

Finite-Sample Metric Non-Collapse for Geometrically Supervised Latent World Models in Control

Alain Bensoussan, Minh-Nhat Phung, Minh-Binh Tran

首次发表
浏览论文内容

中文总结 AI 辅助

该研究为非线性确定性控制系统的学习型世界模型构建三部分数学理论,含逼近理论、有限样本几何保证及确定性规划迁移方法,通过数值实验验证了相关方法的有效性。

中文摘要 AI 辅助

针对非线性确定性控制系统,我们构建了一套三部分的数学理论,用于构建度量保真的学习型世界模型。第一,在逼近理论方面,我们构造了光滑精确的潜实现,并验证了范数约束下的张量积B样条类所需的有限容量C^{1,1}逼近性质,该类具有与容量无关的正则预算。第二,在有限样本几何方面,我们引入了仅编码器的局部-全局度量铰链,其方向项与分离对项可防止无穷小崩溃与全局折叠。尽管预测损失基于观测值,但该惩罚项的评估通过可观测状态距离与切方向利用状态度量监督实现。在正则可观测因子假设、低Ahlfors覆盖度及均匀C^{1,1}预算下,若显式有限样本偏差低于度量裕度阈值,每一个高于可计算单侧正则化阈值的近似经验极小化器均为逐点余利普希茨的,且满足均匀近似受控半共轭估计,其中逼近、统计及训练优化误差相互分离;该步骤所用的L^2到L^∞指数在利普希茨正则性下是尖锐的。第三,对于确定性规划迁移,度量非崩溃诱导出兼容的利普希茨潜代价,而半共轭性给出均匀轨迹与有限时域代价界及优化器迁移保证,学习型潜代价头通过显式兼容误差参与。数值实验采用有界潜范围、存档固定度量样本、受该定理启发的有限容量系数盒样条代理,以及受控摆上的潜模型预测控制。

英文摘要

We establish a finite-sample learning-to-control theory for geometrically supervised latent models of nonlinear deterministic systems. Geometric supervision is used only during training: simulator state, proprioception, or state estimates with independently validated metric and directional error bounds supply observable-state distances and tangent directions, while deployment remains observation- and action-conditioned. We introduce an encoder-only local--global metric hinge that enforces directional resolution and separated-state discrimination. Under regular observable-factor, coverage, finite-capacity approximation, and uniform $C^{1,1}$ hypotheses, a computable one-sided regularization regime has a strong selection property: with high probability, every approximate empirical minimizer is simultaneously pointwise co-Lipschitz and uniformly approximately semiconjugate to the controlled dynamics. Approximation, sampling, and optimization errors remain explicit and separate. Norm-constrained tensor-product B-spline classes constructively realize the approximation hypotheses, and the interpolation exponent converting mean residual control into a uniform bound is sharp. A modular deterministic corollary transfers the learned certificates to trajectory, finite-horizon cost, learned-cost-head, and optimizer guarantees, while a validated finite-net result enables sharper model-specific certification. Controlled experiments isolate collapse and folding, quantify the analytic certificate's reserve, and demonstrate the control benefit of restored metric resolution. The principal contribution is a complete finite-sample implication from approximate empirical optimization to metric faithfulness, uniform controlled dynamics, and reliable planning for the same learned model.

发表机构

  • Naveen Jindal School of Management University of Texas at Dallas, Richardson, TX, 75080 USA
  • Department of Mathematics Texas A\&M University, College Station, TX, 77843 USA
  • Naveen Jindal School of Management University of Texas at Dallas, Richardson, TX 75080, USA
  • Department of Mathematics Texas A\&M University, College Station, TX 77843, USA

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑