AI 中文总结
研究重复两人正规形式博弈中对基于计数学习者的欺骗行为,形式化相关玩法与欺骗奖励概念,建立结构结果,设计算法优化对抗不同情况的对手,证明近似最优欺骗收益针对ERM学习规则是NP难的,并通过实验测量随机博弈中的欺骗奖励。
AI 中文摘要
在重复博弈中,对手常通过观察我们过去的行为来预测下一步行动,这使我们能够欺骗他们。我们研究在与基于计数的学习者进行的重复两人正规形式博弈中的欺骗行为,这类学习者的行为仅取决于我们过去执行每个行动的频率。我们形式化了欺骗性和非欺骗性玩法,引入了欺骗奖励的概念,即最佳欺骗策略相对于最佳固定混合策略的收益增益。我们建立了一般和博弈中欺骗的结构结果。针对行动空间或对手记忆较小的情况,设计了精确的动态规划来优化对抗任何基于计数的学习者;对于对手记忆或时间范围较大的设置,开发了近似算法。还提供了一种用于学习欺骗基于计数学习规则未知的对手的近似算法。此外,我们证明了针对经典经验风险最小化(ERM)学习规则近似最优欺骗收益是NP难的。最后,我们通过实验测量了具有独立同分布收益的随机博弈中的欺骗奖励。
英文摘要
In repeated games, opponents often predict what we'll do next by looking at what we have done so far. This allows us to deceive them: we can deliberately behave one way for a period of time to shape their expectations, then switch strategies to profit from the induced response. We study deception in repeated two-player normal-form games against count-based learners, whose behavior depends only on how often we have played each action in the past. We formalize deceptive and non-deceptive play, and introduce the notion of a deception bonus, the payoff gain of the best deceptive strategy over the best fixed mixed strategy. We establish structural results on deception in general-sum games. We design exact dynamic programs for optimizing against any count-based learner when the action space or opponent's memory is small, and develop approximation algorithms for settings where the opponent's memory or the time horizon is large. We also provide an approximation algorithm for learning to deceive an opponent whose count-based learning rule is unknown. To complement our algorithmic results, we show that approximating the optimal deceptive payoff against the classic Empirical Risk Minimization (ERM) learning rule is NP-hard, including obtaining any constant-factor approximation or even a $T^α$-additive approximation for any $0 < α< 1$. Finally, we empirically measure the deception bonus in random games with i.i.d. payoffs.