发表机构
Zhejiang University; Beijing Institute for General Artificial Intelligence (BIGAI); Shandong University(浙江大学; 北京通用人工智能研究院; 山东大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出领导者-防御者-攻击者博弈的嵌套均衡概念,推导仿射控制律与获胜区域,并用滚动时域算法实现,数值实验验证了预测时域的影响及与独立控制基线的差异。
AI 中文摘要
我们构建了一个离散时间的领导者-防御者-攻击者博弈,该博弈将保护组内的分层交互与对抗外部攻击者的竞争相结合。领导者旨在到达指定需求点同时避免被捕获,而防御者旨在拦截攻击者同时保持在领导者附近。为捕捉他们各自的目标和不对称交互,我们引入了一个嵌套均衡概念,将领导者-防御者Stackelberg关系与攻击者的Nash型最优响应条件相结合。对于线性动力学和二次目标,我们推导了仿射响应律和向后递归,以及唯一阶段解存在的充分条件。我们进一步将到达、捕获和拦截条件表示为初始联合状态的二次不等式,刻画了在所得策略下的获胜区域。一种滚动时域算法利用更新的状态信息实现所计算的控制律。在随机初始配置上的数值实验考察了预测时域的影响,并将所提方法与独立控制基线进行比较,展示了博弈结果和智能体实际成本的差异。
英文摘要
We formulate a discrete-time leader-defender-attacker game that combines hierarchical interaction within a protective group with competition against an external attacker. The leader seeks to reach a prescribed demand point while avoiding capture, whereas the defender seeks to intercept the attacker while remaining near the leader. To capture their distinct objectives and asymmetric interactions, we introduce a nested equilibrium concept combining a leader-defender Stackelberg relationship with a Nash-type best-response condition for the attacker. For linear dynamics and quadratic objectives, we derive affine response laws and a backward recursion, together with sufficient conditions for unique stagewise solutions. We further express arrival, capture, and interception conditions as quadratic inequalities in the initial joint state, characterizing winning regions under the resulting policy. A receding-horizon algorithm implements the computed control laws using updated state information. Numerical experiments over randomized initial configurations examine the influence of the prediction horizon and compare the proposed method with an independent-control baseline, illustrating differences in game outcomes and the agents' realized costs.