发表机构
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究在线控制方法,通过反事实跟踪,在已知线性动态系统受干扰及凸成本情况下,模拟基准策略形成参考并跟踪,建立PAC - Bayes遗憾保证,与稳定线性动态控制器系统级响应球竞争,给出最优依赖性及匹配下界。
AI 中文摘要
我们开发了一种在线控制方法,它能与一般因果策略类别竞争,超越大多数现有算法所使用的线性控制器类别。在\(T\)轮的时间范围内,考虑一个已知的线性动态系统,该系统受到对抗性干扰且每次行动后会揭示凸成本。该方法在已揭示的历史上模拟基准策略,利用它们的反事实状态 - 输入对形成移动参考,并应用固定稳定控制器在物理系统上跟踪该参考,即反事实跟踪。此方法适用于任何可从已揭示历史模拟且其反事实状态 - 输入对在每轮具有有界直径的可测因果策略类别。我们建立了PAC - Bayes遗憾保证,该保证对策略的每个后验都成立,并取决于其相对于所选先验的相对熵。在具有有界脉冲响应增益跟踪器的固定装置上,当\(\log N = O(T)\)时,有限的\(N\)个策略类别允许具有\(\sqrt{T\log N}\)的极小极大最优\(T\)和\(N\)依赖性。作为核心应用,我们与稳定线性动态控制器的系统级响应球竞争。该球界定了与固定跟踪器的总脉冲响应偏差,但不施加共同的衰减包络、记忆长度或控制器阶数限制。据我们所知,这是第一个对此类进行统一的在线控制保证。匹配的下界表明我们的保证在常数范围内是紧的。
英文摘要
We study online control of a known linear dynamical system with adversarial costs and bounded disturbances, measuring regret against a general class of benchmark policies. We introduce counterfactual tracking, which separates the challenge of learning from the challenge of controlling the system. An online learner builds a reference trajectory by selecting or averaging the trajectories that the benchmark policies would have generated under the realized costs and disturbances, and a corrective law steers the system toward that reference. Charging each change in the reference its recovery cost (the cost of steering the system onto the new reference) reduces the problem to online learning with switching costs. Conversely, under additional natural assumptions, we show that this reduction is tight: the two problems have the same minimax regret up to a system-dependent factor, uniformly over horizons and policy classes. The reduction gives sharp regret guarantees for policy classes beyond standard finite-memory parameterizations. For a class of $N$ possibly nonlinear or history-dependent policies, it achieves $O(\sqrt{T\log N})$ regret over $T$ rounds, provided their trajectories remain within a bounded distance of one another and recovery costs are bounded. For the full $\ell_1$ ball of disturbance-response controllers, it achieves $O(\sqrt{T\log T})$ regret, which is minimax optimal in $T$, without assuming a common decay rate for disturbance effects. The framework also improves the best known regret bounds for linear state-feedback policies.