AI 中文总结
本文针对隐藏拜占庭攻击下的多智能体系统在线协同控制问题,分析攻击者信息对模型几何特性的影响,推导安全学习的信息论极限,提出鲁棒学习者并给出遗憾界,为可靠多智能体系统提供理论与算法基础。
AI 中文摘要
我们研究拜占庭攻击下多智能体系统的在线协同控制问题。具体而言,存在一个未知的固定智能体子集为拜占庭智能体,它们在观测到团队规划的联合行动后,可隐秘地覆盖自身坐标。学习者可观测到规划行动、公共奖励和公共状态,但无法观测到覆盖行为或执行的联合行动。我们的目标是安全性:针对最坏情况的覆盖行为优化团队性能,以达到最优安全值。我们首先证明攻击者的信息决定了问题的几何特性:观测到规划行动的攻击者会诱导出一个精确的(s,a)-矩形鲁棒马尔可夫决策过程(MDP),其行是覆盖诱导的公共结果律的凸包;而未观测到规划行动的盲攻击者则会诱导出一个s-矩形模型。随后,我们确定了安全学习的信息论极限,证明安全遗憾可精确分解为针对生成数据的响应的回报遗憾与累积响应差距D_K之和。两个无法区分的单步实例会导致Ω(K)的期望安全遗憾,而回报遗憾为零,这表明对D_K的依赖不可避免。最后,我们提出了一种与阶段绑定的鲁棒估计-决策学习者,并证明其遗憾界为Õ(H²S√(AK)) + E[D_K]。本研究为拜占庭攻击下可靠多智能体系统提供了全面的理论与算法基础。
英文摘要
We study online cooperative control of a multi-agent system under Byzantine attacks. Namely, an unknown, fixed subset of agents are Byzantine comprised and can stealthily overwrite its own coordinates of the team's planned joint action after observing that plan. The learner observes planned actions, public rewards, and public states, but neither the overwrite nor the executed joint action. Our objective is security: to optimize the team performance against the worst overwrites and achieve the optimal security value. We first show that the attacker's information determines the geometry. An attacker that observes the planned action induces an exact $(s,a)$-rectangular robust Markov decision process (MDP) whose rows are convex hulls of overwrite-induced public-outcome laws, whereas a blind attacker induces an $s$-rectangular model. We then identify the information-theoretic limit of security learning, showing that the security regret decomposes exactly into return regret against the response generating the data and a cumulative response gap $D_K$. Two indistinguishable horizon-one instances force $Ω(K)$ expected security regret while return regret is zero, showing that dependence on $D_K$ is unavoidable. Finally, we develop a stage-tied robust estimation-to-decisions learner and prove a regret bound of $\widetilde{\mathcal O}\!\left(H^2S\sqrt{AK}\right)+\mathbb E[D_K]$. Our studies thus provide comprehensive theoretical and algorithmic foundations of reliable multi-agent systems under Byzantine attacks.
Commentspreprint; Work in progress