arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.13651eess.SYcs.LGcs.SYmath.OC

一致模型追踪是极小极大最优的:大参数不确定性下标量对抗自适应控制的精确值

Consistent Model Chasing Is Minimax Optimal: The Exact Value of Scalar Adversarial Adaptive Control under Large Parametric Uncertainty

Dimitar Ho

首次发表
浏览论文内容

中文总结 AI 辅助

该研究精确求解了大参数不确定性下对标量系统的对抗自适应控制问题,得出最优值为1+Δ,证明了基于一致模型追踪的中点等价确定性无差拍控制是极小极大最优的,且学习完全被动。

中文摘要 AI 辅助

我们精确求解了一个对抗扰动下自适应控制的基础问题:对标量系统$x_{t+1} = ax_t + u_t + w_t$($x_0=0$,$\\|w\\|_\n\le 1$)进行调节,其中恒定极点$a \in [-Δ, Δ]$的符号和幅值均未知,且$Δ$可任意大。尽管该系统看似简单,但据我们所知,在该准则下,任何具有任意大小参数不确定性的自适应控制问题中,因果控制器针对对抗对$(a, w)$所能保证的最小最坏峰值$\\|x\\|_\n$(即该博弈的值)从未被确定;现有理论提供的是稳定性证明、增益界和悔界速率,而非该精确值。\n对任意$Δ>0$,该值为$γ^\star(Δ) = 1 + Δ$。其中加项1是扰动的不可约代价,$Δ$是单次不可避免的辨识尖峰的精确代价。最优策略是在集合隶属一致区间的中点处执行等价确定性无差拍控制,这是鲁棒预言机×一致模型追踪架构的一个实例。该架构是强制的,而非仅仅是充分的:定义$θ_t := -u_t/x_t$可将每个因果控制器表示为预言机-选择器的组合,而最优性要求选择器在关键历史时刻取中点。\n经典和现代的标准工具均存在可量化的缺陷:探测在产生收益前就会受到惩罚;一旦需要自适应,在亚扰动激励下的承诺是致命的;乐观主义退化为平局打破,或渐近付出至少两倍于最优值的代价;而悔界证明对双向的最坏峰值均不敏感。最优律不包含任何探索机制,其学习完全是被动的。这些结果首次为一致模型追踪作为对抗自适应控制的设计原则提供了精确的最优性证明。

英文摘要

We solve exactly a fundamental problem of adaptive control against adversarial disturbances: regulate the scalar system $x_{t+1} = ax_t + u_t + w_t$, $x_0=0$, $\|w\|_\infty \le 1$, where the constant pole $a \in [-Δ, Δ]$ is unknown in sign and magnitude and $Δ$ is arbitrarily large. Elementary as the system looks, the least worst-case peak $\|x\|_\infty$ that a causal controller can guarantee against an adversarial pair $(a, w)$ (the value of this game) has, to our knowledge, never been determined for any adaptive control problem with parametric uncertainty of arbitrary size under this criterion; existing theory supplies stability certificates, gain bounds, and regret rates, not the value. That value is $γ^\star(Δ) = 1 + Δ$ for every $Δ>0$. The summand $1$ is the irreducible price of the disturbance, and $Δ$ the exact price of a single, unavoidable identification spike. The optimal policy is certainty-equivalent deadbeat control at the midpoint of the set-membership consistent interval, an instance of the robust oracle $\times$ consistent model chasing architecture. The architecture is forced, not merely sufficient: writing $θ_t := -u_t/x_t$ exhibits every causal controller as an oracle-selector composition, and optimality pins the selector to the midpoint at the critical histories. The standard tools, classical and modern, each fail quantifiably: probing is punished before it pays, commitment is fatal at sub-disturbance excitation once adaptation is necessary, optimism degenerates to tie-breaking or pays asymptotically at least twice the optimum, and regret certificates are blind to the worst-case peak in both directions. The optimal law contains no exploration mechanism, its learning purely passive. These results give the first exact optimality certificate for consistent model chasing as a design principle for adversarial adaptive control.

补充信息

↑