发表机构
University of California, San Diego; University at Buffalo, State University of New York(加利福尼亚大学圣迭戈分校; 纽约州立大学布法罗分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究多智能体学习中探索的隐私脆弱性,证明欺骗者可通过耦合探索信息将学习动态导向欺骗性纳什均衡,并保持最优收敛速率。
AI 中文摘要
随机化探索是老虎机学习、多智能体强化学习和零阶策略搜索的核心,然而其独立性和隐私性通常仅被视为技术假设。我们表明这些属性对安全目的至关重要,并演示了一个对抗性智能体如何利用关于另一个智能体探索的特权信息。我们在最小两人强单调设置中分析了一个欺骗者-受害者对,其中欺骗性玩家获得与受害者探索仅相关的泄漏信号。我们表明,通过将自己的探索行动与该信息耦合,欺骗性玩家注入了一种外部性,将学习动态引导到一个新的稳态,称为欺骗性纳什均衡(DNE)。我们证明欺骗性老虎机学习(DBL)动态收敛到DNE的任意小邻域,同时保持最优收敛速率。有趣的是,我们的分析在放宽标准老虎机优化文献中的二阶光滑性条件的同时达到了这些最优速率。我们刻画了欺骗严格改变稳态的条件及其对欺骗者成本的影响,并在资源分配博弈中说明了结果。
英文摘要
Randomized exploration is central to bandit learning, multi-agent reinforcement learning, and zeroth-order policy search, yet its independence and privacy are usually only treated as technical assumptions. We show that these properties are critical for security purposes and demonstrate how an adversarial agent can exploit privileged information on another agent's exploration. We analyze a deceiver-victim pair in the minimal two-player strongly monotone setting, where a deceptive player obtains leaked signals that are merely correlated with the victim's exploration. We show that, by coupling their own exploratory action with this information, the deceptive player injects an externality that steers the learning dynamics to a new steady state, called the deceptive Nash equilibrium (DNE). We prove that the deceptive bandit learning (DBL) dynamics converge to an arbitrarily small neighborhood of the DNE while retaining optimal convergence rates. Interestingly, our analysis attains these optimal rates while relaxing second-order smoothness conditions from standard bandit optimization literature. We characterize conditions under which deception strictly shifts the steady state and its effect on the deceiver's cost, illustrating the results in a resource-allocation game.