发表机构
University of Waterloo(滑铁卢大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对偏微分方程提出目标无关控制框架,基于联合嵌入预测架构,通过离线训练小型二维ViT编码器等,经模型预测路径积分控制器复用。在相关基准测试中,KE探测器规划提升奖励、降低误差,验证了潜在动力学与目标无关及校准可观测量用于状态控制的优势。
AI 中文摘要
我们提出了一个围绕联合嵌入预测架构(JEPA)构建的偏微分方程(PDE)目标无关控制框架。小型二维视觉Transformer(ViT)编码器和动作条件潜在动力学在离线时无奖励或下游目标的情况下进行训练,冻结后由模型预测路径积分(MPPI)控制器复用。我们发现,当有可用的控制目标时,将其应用于明确的物理可观测量(假设具有单射性)比在学习的潜在空间中最小化原始欧几里得距离($L^2$)更好。对于冻结潜在展开上的学习线性动能(KE)探测器,我们可以以$R^2 = 0.989$再现保留轨迹,同时无需更改基础世界模型。在PDE控制健身房二维纳维 - 斯托克斯基准测试中,使用KE探测器规划将匹配的50集原生奖励从潜在 - $L^2$规划的$-12.08 \pm 0.86$提高到$-10.90 \pm 0.91$(95%置信区间),同时将最后四分之一速度场均方根误差从$0.0765$降低到$0. 0692$。在三个故意保留的、不同的、非周期性目标上,KE规划相对于潜在 - $L^2$规划将后期场均方根误差降低了53%($0.0220$对$0.0469$),赢得了所有30对剧集。相同的冻结模型还支持通过直接调节KE实现围绕稳定配置的稳定控制,平均相对误差为$2.7\%$。虽然潜在探测器对测量噪声和缺失像素很脆弱,但我们相信这些结果支持这样的观点,即潜在动力学可以保持动态和目标无关,而校准的可观测量(假设它们保证唯一延续)可能是状态控制的更好目标。
英文摘要
We present a goal-agnostic control framework for partial differential equations (PDEs) built around an end-to-end joint-embedding predictive architecture (JEPA). A lightweight 2D vision-transformer (ViT) and action-conditioned latent dynamics are trained offline without a reward or downstream goal, before being frozen and reused by a model-predictive path integral (MPPI) controller. We minimize a control objective in the latent space, initially expressed via the $L^2$ distance and additionally illustrate the benefit of recasting the control objective in terms of an explicit physical observable when available. By instead minimizing the tracking error for a learned linear kinetic-energy (KE) probe on the frozen latent-state rollouts, we demonstrate the ability to reproduce the control of held-out trajectories with $R^2=0.989$, while requiring no change to the underlying world model. For a controlled 2D Navier--Stokes benchmark, using a KE-probe within MPPI planning improves the mean native reward from $-12.08\pm0.86$ for latent-$L^2$ tracking to $-10.90\pm0.91$ (95\% CI), all while lowering last-quarter velocity-field RMSE from $0.0765$ to $0.0692$. Across three intentionally withheld, dissimilar, aperiodic targets, KE planning lowers late field RMSE by $53\%$ relative to latent-$L^2$ planning ($0.0220$ versus $0.0469$), winning across 30 paired comparisons. The same frozen model also supports stabilization around a steady-state configuration via direct regulation of KE, achieving $2.7\%$ mean relative error. While the latent probe proves brittle to measurement noise and missing pixels, our findings support the claim that latent dynamics can remain flexible and goal-agnostic, particularly when calibrated observables (granted they guarantee unique continuation) are a suitable objective for state control.
CommentsAssociated code will be open sourced alongside a later, updated submission