发表机构
University of Michigan; Carnegie Mellon University(密歇根大学; 卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种基于Transformer的反馈策略,通过模仿LQR控制实现泛化,并给出有限时域闭环近最优性的概率证书,在28个基准系统上验证了低违反概率。
AI 中文摘要
本文为基于Transformer的反馈策略建立了闭环性能证书。该策略经过训练,以模仿跨一系列异构多输入多输出(MIMO)线性时不变(LTI)系统族的最优线性二次调节器(LQR)控制。首先,我们为训练期间最小化的模仿损失建立了有限样本超额风险界。其次,对于每个固定的问题实例,我们推导了区域闭环保证,包括前向不变运行区域和相对于最优轨迹偏差的最坏情况界。我们的主要结果是有限时域闭环近最优性的概率证书。利用精确的LQR成本恒等式,我们将超额成本表示为可测量的每轨迹统计量,并使用独立的校准和验证轨迹来获得其违反概率的高置信界。我们在$28$个基准系统上评估该证书。这使用基础策略用于已见系统,并在未见系统上使用系统特定的微调副本,每次轨迹从相应的认证分布中抽取被控对象、成本和初始条件。所有每个系统的证书的违反概率均低于$3.1\%$,每个置信度为$95\%$;二十个系统的次优性低于$10\%$,最紧的阈值等于$4.8\ imes10^{-6}$。
英文摘要
This letter develops closed-loop performance certificates for a transformer-based feedback policy. The policy is trained to imitate optimal Linear Quadratic Regulator (LQR) control across a family of heterogeneous Multiple-Input, Multiple-Output (MIMO) Linear Time-Invariant (LTI) systems. First, we establish a finite-sample excess-risk bound for the imitation loss minimized during training. Second, for each fixed problem instance, we derive regional closed-loop guarantees consisting of a forward-invariant operating region and a worst-case bound on deviation from the optimal rollout. Our main result is a probabilistic certificate for finite-horizon closed-loop near-optimality. Using an exact LQR cost identity, we express excess cost as a measurable per-rollout statistic and use independent calibration and validation rollouts to obtain a high-confidence bound on its violation probability. We evaluate the certificate on $28$ benchmark systems. This uses the base policy on seen systems and system-specific fine-tuned copies on unseen systems, with each rollout drawing the plant, cost, and initial condition from the corresponding certification distribution. All per-system certificates have violation probabilities below $3.1\%$, each at $95\%$ confidence; twenty systems certify suboptimality below $10\%$, with the tightest threshold equal to $4.8\times10^{-6}$.