发表机构
Singapore University of Technology and Design(新加坡科技设计大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究连续博弈中仅依赖成对偏好反馈的学习问题,提出序数遗憾基准与块归一化伪梯度动力学,实现次线性遗憾及纳什收敛,无需重建基数效用。
AI 中文摘要
我们研究了当玩家仅接收成对偏好反馈时连续博弈中的学习问题,该反馈揭示了两个动作中哪一个更受偏好,但既不提供收益值也不提供偏好强度。我们首先证明,标准的外部遗憾和粗相关均衡(CCE)无法从这种序数信息中识别:相同的博弈序列可以在两个序数等价的博弈中产生零遗憾和线性遗憾,而所有与相同偏好一致的基本基数表示中保持CCE的分布恰好是那些支持纯纳什均衡的分布。受这一差距的启发,我们发展了一种基于归一化单边偏好方向的一阶序数理论,引入了序数遗憾基准和相应的均衡概念。我们证明了块归一化伪梯度动力学实现了次线性序数遗憾,并在额外结构下保证了纳什收敛性。然后,我们使用单比较估计器从有限的成对比较中实现这些动力学。每轮每位玩家进行一次比较,所得到的算法针对任意对手行为实现了次线性有限分辨率序数遗憾,并且在序数势博弈中,几乎必然地最后迭代收敛到纳什集。我们的结果直接提供了来自偏好反馈的遗憾、动力学和均衡保证,而无需重建基数效用。
英文摘要
We study learning in continuous games when players receive only pairwise preference feedback, revealing which of two actions is preferred but neither payoff values nor preference magnitudes. We first show that standard external regret and coarse correlated equilibria (CCE) are not identifiable from this ordinal information: the same sequence of play can incur zero and linear regret in two ordinally equivalent games, while the distributions that remain CCE across all cardinal representations consistent with the same preferences are exactly those supported on pure Nash equilibria. Motivated by this gap, we develop a first-order ordinal theory based on normalized unilateral preference directions, introducing an ordinal directional regret benchmark and corresponding equilibrium notions. We show that block-normalized pseudogradient dynamics achieve sublinear ordinal regret and, under additional structure, Nash-convergence guarantees. We then use a single-comparison estimator to implement these dynamics from finite pairwise comparisons. With one comparison per player and round, the resulting algorithm achieves sublinear finite-resolution ordinal regret against arbitrary opponent behavior and, in ordinal potential games, almost-sure last-iterate convergence to the Nash set. Our results provide regret, dynamics, and equilibrium guarantees directly from preference feedback without reconstructing cardinal utilities.
Comments44 pages, 1 figure