发表机构
Boston University(波士顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究安全自适应控制中安全约束是否允许有信息实验,定义预承诺信息为承诺前学习者可见定律间KL散度,通过因果约简得出有界预承诺信息会留预言机差距,在约束线性系统中建立障碍并证明特殊情况恢复及推导相关证书。
AI 中文摘要
安全自适应控制是在学习轨迹本身的安全保证下进行在线适应。控制器可使用任何因果的、依赖历史的规则,并随数据到达在不同环境中采取不同行动。只有其安全保证是统一的:同一规则在每个初始似然模型下都必须满足。性能是相对于知道已实现模型的安全预言机来衡量的。许多有限时间分析假设均匀安全闭环的持续激励,所以数据能区分每对需要不同控制决策的模型。在此假设下,可行性已解决,只剩速率问题。我们转而问:安全约束是否允许这样一个有信息的实验?虽然一种替代情况仍似合理,但控制器必须在其下保持安全延续。我们称排除这种延续的第一个行动为承诺。机会安全仅允许在替代情况下罕见事件上进行承诺,且证据必须提前到达:承诺行动产生的观察太晚。我们将预承诺信息定义为在承诺前停止的学习者可见定律之间的KL散度。我们的主要结果是一个因果约简。承诺规则决定(1)在替代情况下安全允许承诺的概率,(2)保持不承诺的目标侧成本,(3)以及决策时可用的信息。有界预承诺信息因此不可避免地留下固定比例的预言机差距。如果差距是\({\Omega}(T)\),每个均匀安全策略都有线性遗憾。我们在具有二次调节成本的约束线性系统中建立了障碍。我们还证明了在特殊情况下的恢复,并为确定性线性高斯系统推导了半定上证书。
英文摘要
Safe adaptive control is online adaptation under a safety guarantee on the learning trajectory itself. The controller may use any causal, history-dependent rule and act differently across environments as data arrive. Only its safety guarantee is uniform: the same rule must satisfy it under every initially plausible model. Performance is measured against a safe oracle that knows the realized model. Many finite-time analyses assume persistent excitation of the uniformly safe closed loop, so the data distinguish every pair of models requiring different control decisions. Under that assumption, feasibility is already settled; only the rate remains. We ask instead: Do the safety constraints permit such an informative experiment at all? While an alternative remains plausible, the controller must preserve a safe continuation under it. We call the first action that forecloses such a continuation commitment. Chance safety allows commitment only on an event rare under the alternative, and the evidence must arrive beforehand: the observation generated by the committing action is too late. We define precommitment information as the KL divergence between learner-visible laws stopped before commitment. Our main result is a causal reduction. The commitment rule determines (1) the probability that safety permits commitment under the alternative, (2) the target-side cost of remaining noncommittal, (3) and the information available when the decision is made. Bounded precommitment information therefore leaves a fixed fraction of the oracle gap unavoidable. If the gap is Ω(T), every uniformly safe policy has linear regret. We establish the obstruction in a constrained linear system with quadratic regulation cost. We also prove recovery in special cases and derive semidefinite upper certificates for deterministic linear-Gaussian systems.