逆博弈中的学习:具有概率保证的可处理训练
Learning in Inverse Games: Tractable Training with Probabilistic Guarantees
浏览论文内容
中文总结 AI 辅助
本文提出一种可处理的凸优化框架,通过一阶和Nikaido-Isoda损失学习静态与动态非合作博弈,利用RKHS中的Helmholtz-Hodge分解确保强单调性,并给出有限样本预测保证,实验验证了预测精度与噪声鲁棒性。
中文摘要 AI 辅助
逆博弈论旨在从观察到的均衡行为中学习智能体的未知目标。现有的基于残差的方法可能导致非凸问题,并且不能确保所学博弈的强单调性,从而限制了可靠的均衡预测。我们开发了一个可处理的凸框架,用于学习静态和动态非合作博弈,采用一阶和Nikaido-Isoda(NI)损失。专门的算子参数化,包括在再生核希尔伯特空间(RKHS)中的Helmholtz-Hodge分解,强制了二次、非参数、非二次和线性二次动态博弈的强单调性。我们利用Rademacher复杂度建立了有限样本的样本外预测保证,与非二次设置中的现有界相匹配。对于含噪声的序列数据,我们开发了一种鲁棒的滚动时域学习方案,其自适应正则化具有最大后验(MAP)解释,并通过随机迹估计进行计算。数值实验,包括一个风格化的自动驾驶车辆碰撞避免应用,展示了预测准确性和对测量噪声的鲁棒性。
英文摘要
Inverse game theory seeks to learn agents' unknown objectives from observed equilibrium behavior. Existing residual-based approaches can lead to non-convex problems and need not ensure strong monotonicity of the learned game, limiting reliable equilibrium prediction. We develop a tractable convex framework for learning static and dynamic non-cooperative games using first-order and Nikaido-Isoda (NI) loss. Specialized operator parameterizations, including a Helmholtz--Hodge decomposition in a reproducing kernel Hilbert space (RKHS), enforce strong monotonicity for quadratic, nonparametric, non-quadratic, and linear-quadratic dynamic games. We establish finite-sample out-of-sample prediction guarantees, using Rademacher complexity to match existing bounds for the non-quadratic setting. For noisy sequential data, we develop a robust receding-horizon learning scheme whose adaptive regularization admits a maximum a posteriori (MAP) interpretation and is computed using randomized trace estimation. Numerical experiments, including a stylized autonomous-vehicle collision-avoidance application, demonstrate predictive accuracy and robustness to measurement noise.
发表机构
- Delft University of Technology(代尔夫特理工大学)
- Erasmus University(伊拉斯姆斯大学)
- University of Toronto(多伦多大学)
机构由 AI 辅助整理,请以论文原文为准。