arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.18067cs.GT

在秩-1博弈中针对学习的规划

Planning Against Learning in Rank-1 Games

William Overman

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对满足$\text{rank}(A+B)=1$的双矩阵博弈,证明其纳什均衡可多项式时间计算,但针对学习动态的最优规划是NP困难的,区分了高效均衡计算与策略规划。

中文摘要 AI 辅助

学习算法常被用于重复多智能体环境中的决策。当另一玩家了解学习者如何根据过往经验调整时,该玩家可在多轮中进行策略性规划,以影响学习者的未来行为。近期研究表明,针对复制动态(Replicator Dynamics,乘法权重更新的连续时间类似物)进行优化在零和博弈中是易处理的,但在无限制的一般和博弈中可能很困难。我们研究了零和之外的首个结构化类别:满足$\text{rank}(A+B)=1$的双矩阵博弈,对于这类博弈,纳什均衡可在多项式时间内计算。我们的主要结果显示,这种均衡易处理性无法扩展到针对学习动态的规划。除非$\text{P}=\text{NP}$,即使$\text{rank}(A+B)=1$、学习者从均匀状态开始且优化器被限制为常数策略,在固定加性常数内近似优化器的最优连续时间回报仍是NP困难的。这种困难性对于有界支付矩阵和多项式有界时间范围依然存在。我们通过几个易处理特殊情况的结构表征补充了这一结果。因此,秩-1博弈已将高效均衡计算与针对学习对手的策略性规划区分开来。

英文摘要

Learning algorithms are often used to make decisions in repeated multi-agent environments. When another player understands how a learner adapts from past experience, that player can plan strategically across rounds to influence the learner's future behavior. Recent work shows that optimizing against Replicator Dynamics, the continuous-time analogue of Multiplicative Weights Update, is tractable in zero-sum games but can be hard in unrestricted general-sum games. We study the first structured class beyond zero sum: bimatrix games satisfying $\text{rank}(A+B)=1$, for which Nash equilibria can be computed in polynomial time. Our main result shows that this equilibrium tractability does not extend to planning against learning dynamics. Unless $\mathsf{P}=\mathsf{NP}$, approximating the optimizer's optimal continuous-time reward within a fixed additive constant is NP-hard even when $\text{rank}(A+B)=1$, the learner starts from the uniform state, and the optimizer is restricted to constant strategies. The hardness persists for bounded payoff matrices and polynomially bounded horizons. We complement this result with structural characterizations of several tractable special cases. Thus rank-one games already separate efficient equilibrium computation from strategic planning against a learning opponent.

↑