逆滚动时域线性二次调节器问题的可辨识性与泛化性刻画
Characterizing Identifiability and Generalization for Inverse Receding-Horizon Linear-Quadratic Regulator Problems
浏览论文内容
中文总结 AI 辅助
本文研究滚动时域LQR中目标推断的可辨识性,刻画唯一可辨识条件,并证明仿射包内动作一致、包外有误差上界,数值验证噪声下预测有效。
中文摘要 AI 辅助
我们考虑滚动时域线性二次调节器(LQR)背景下目标推断的问题。在此设定中,我们获得序列的状态-动作观测,其中每个观测到的动作是新求解的有限时域LQR问题的第一个控制量。我们刻画了该问题的目标何时能从这些观测中唯一可辨识,以及何时额外的观测不提供关于目标的任何新信息。然后,我们分析在未见状态下的动作预测,并表明所有复现观测动作的目标在观测状态仿射包内产生相同的动作;在该包之外,我们推导出预测误差的上界。此外,我们证明,当仅有线性目标项未知时,在每个状态下精确预测均成立。最后,数值结果表明,即使在随机观测噪声下,重新优化推断出的目标也能在不同规划时域下对未见状态实现准确的动作预测。
英文摘要
We consider the problem of objective inference in the context of receding-horizon linear-quadratic regulator (LQR). In this setting, we are given sequential state-action observations, where each observed action is the first control of a newly solved finite-horizon LQR problem. We characterize when the objective of that problem is uniquely identifiable from these observations and when additional observations provide no new information about the objective. We then analyze action prediction at unseen states and show that all objectives reproducing the observed actions yield identical actions throughout the affine hull of the observed states; outside this hull, we derive an upper bound on the prediction error. % and deriving a prediction-error bound outside this hull. Additionally, we show that, when only the linear objective terms are unknown, exact prediction holds at every state. Finally, numerical results show that, even under stochastic observation noise, re-optimizing an inferred objective enables accurate action prediction at unseen states across different planning horizons.
发表机构
- University of Wisconsin–Madison(威斯康星大学麦迪逊分校)
- The University of Texas at Austin(德克萨斯大学奥斯汀分校)
机构由 AI 辅助整理,请以论文原文为准。