arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

关于效率脆弱性:存在偏离者时的非耦合学习

On the Fragility of Efficiency: Uncoupled Learning with a Deviant

Vade Shah, Jason R. Marden

arXiv 2610.06548首次发表:更新:

发表机构

University of California, Santa Barbara(加州大学圣塔芭芭拉分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究探讨完全非耦合福利最大化学习算法对偏离玩家的脆弱性,证明单个偏离者(固定行动、随机或策略性)可显著破坏长期效率,且任何此类算法均无法在所有博弈中鲁棒,表明鲁棒性需沟通或协调。

AI 中文摘要

博弈学习中的一个重要问题是,玩家是否能够在没有沟通或协调的情况下,学会实现高效结果,即最大化其效用总和的结果。完全非耦合的学习算法,仅使用每个玩家自身过去的行动和效用,可以在广泛的博弈类别中保证这一点,但通常要求每个玩家都遵循相同的算法。在本工作中,我们探讨能否设计对偏离玩家具有鲁棒性的福利最大化学习算法。为研究此问题,我们以一个被广泛研究的福利最大化学习算法作为测试平台,并考虑当除一名玩家外的所有玩家都遵循该算法时会发生什么。首先,我们考虑一个偏离玩家,其要么(i)始终采取相同行动,要么(ii)均匀随机选择其行动。我们证明,这样的玩家可以显著改变该算法的长期行为,导致远离福利最大化的状态,并且没有完全非耦合的学习算法能在每个博弈中对这样的玩家具有鲁棒性。接下来,我们考虑一个按自身利益行事的策略性偏离玩家,并证明该玩家可以采用不同的学习算法来使长期结果偏向自身。此外,我们证明,在任何完全非耦合的福利最大化学习算法下,都存在一个博弈,其中某个玩家通过切换到另一种完全非耦合的学习算法而获益。这些结果表明,完全非耦合的福利最大化学习是脆弱的,即使面对单个偏离玩家也是如此,这暗示鲁棒性可能需要沟通或协调。

英文摘要

An important question in learning in games is whether players can learn to achieve efficient outcomes, those that maximize the sum of their utilities, without communication or coordination. Completely uncoupled learning algorithms, which use only each player's own past actions and utilities, can guarantee this in broad classes of games, but they typically require every player to follow the same algorithm. In this work, we ask whether one can design welfare-maximizing learning algorithms that are robust to a deviant player. To study this question, we take a well-studied welfare-maximizing learning algorithm as a testbed and consider what happens when every player except one follows it. First, we consider a deviant player who either (i) always plays the same action or (ii) selects their action uniformly at random. We show that such a player can significantly alter the long-run behavior of this algorithm, leading to states that are far from welfare-maximizing, and that no completely uncoupled learning algorithm is robust to such a player in every game. Next, we consider a strategic deviant player who acts in their own interest and show that this player can adopt a different learning algorithm to shift the long-run outcome in their favor. Moreover, we show that under any completely uncoupled welfare-maximizing learning algorithm, there is a game in which some player benefits from switching to a different completely uncoupled learning algorithm. These results show that completely uncoupled welfare-maximizing learning is fragile, even to a single deviant player, suggesting that robustness may require communication or coordination.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑