连续时间自适应控制中的尖锐临界极小极大定律与无学习阈值
Sharp Critical Minimax Laws and No-Learning Thresholds in Continuous-Time Adaptive Control
浏览论文内容
中文总结 AI 辅助
研究连续时间自适应控制中未知增益的极小极大遗憾,证明临界速率定律并揭示无学习相变阈值,通过软阈值反馈达到最优。
中文摘要 AI 辅助
我们研究具有未知向量控制增益、标量状态、二次行动成本和光滑凸终端成本的逐回合连续时间控制问题。在标量高斯实验中,令 $\Delta(H)$ 表示相对于零控制的极小极大改进,并设 $\delta=\sqrt{2H^2-1}$。我们证明了临界定律 $$ \Delta(H)\asymp \delta^4\sqrt{\log(1/\delta)} \qquad (\delta\downarrow0), $$ 并给出了匹配的下界和上界。下界源于一个对具有无界振幅的完全自适应控制有效的均匀亏缺-能量不等式,而一个移动软阈值反馈达到了该速率。对于局部参数 $\theta=N^{-1/4}h$,$|h|\le H$,我们证明了在 $N$ 个回合上的归一化极小极大遗憾与一个固定时域高斯序贯控制问题相差在 $O(N^{-1/2})$ 以内,且没有额外的维度因子。终端任务仅通过 $c_g=\operatorname{Var}(g(Z))$ 进入极限。该高斯问题表现出精确的无学习相变: $$ C_d^T(H)=TH^2/2 \iff H^4T\le d^2/2, $$ 从而得到渐近任务边界 $c_gH^4=d^2/2$。
英文摘要
We study episodic continuous-time control with an unknown vector control gain, scalar state, quadratic action cost, and smooth convex terminal cost. In the scalar Gaussian experiment, let $Δ(H)$ denote the minimax improvement over zero control and set $δ=\sqrt2H^2-1$. We prove the critical law $$ Δ(H)\asymp δ^4\sqrt{\log(1/δ)} \qquad (δ\downarrow0), $$ with a matching lower and upper bound. The lower bound follows from a uniform deficit--energy inequality valid for fully adaptive controls with unbounded amplitudes, while a moving soft-threshold feedback attains the rate. For local parameters $θ=N^{-1/4}h$, $|h|\le H$, we show that the normalized minimax regret over $N$ episodes is within $O(N^{-1/2})$ of a fixed-horizon Gaussian sequential control problem, without an additional dimension factor. The terminal task enters the limit only through $c_g=\operatorname{Var}(g(Z))$. The Gaussian problem exhibits an exact no-learning phase transition: $$ C_d^T(H)=TH^2/2 \iff H^4T\le d^2/2, $$ yielding the asymptotic task boundary $c_gH^4=d^2/2$.