arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

正则化自博弈中的均衡选择引导:基于参考策略

Steering Equilibrium Selection in Regularized Self-Play via the Reference Policy

Luis Leal

arXiv 2609.19820首次发表:更新:

AI 中文总结

本研究证明通过将正则化自博弈的参考策略锚定在目标均衡成员并细化,可精确引导均衡选择,实现低误差与低可利用性,并将KL锚重新定义为选择旋钮。

AI 中文摘要

正则化自博弈——DeepNash的Stratego玩法背后的方法家族——通过最佳响应一个缓慢移动的熵正则化参考策略$\rho$,将两人零和策略驱动至纳什均衡。当博弈具有值等价均衡的多面体时,正则化器会静默地打破平局:使用均匀参考时,它选择最大熵成员,即$\rho$在纳什集上的I-投影。参考能否被有目的地用于选择均衡?在五个精确可解博弈加上一个二维多面体上,使用精确最佳响应和独立种子的等价性检验,将参考锚定在目标成员并细化,可将自博弈引导至该成员,平均坐标误差为0.007,中位可利用性为$5\times10^{-5}$,在$\pm0.05$范围内与请求TOST等价;锚定在细化过程中持续存在,并跟随参考而非初始化。选择遵循可达加权I-投影(斜率0.969 [0.950, 0.987])。我们同样强调故事失效之处:固定的离流形参考代价为0.08-0.25可利用性;刚性或平坦族需要更小的镜像步长,由预注册规则设定;边界目标欠冲;曲率预测边界饱和发生的位置(秩相关0.90,p=0.037),而内部精度与曲率无关。表格和MLP引导映射在每个目标处等价于$\pm0.03$(30个种子);匹配对照组显示注意力的稳健特征是多余的种子方差,任何系统性偏移上限为0.018且不显著。相对于最佳响应,选择-稳健性权衡是退化的:引导仅对固定的非均衡对手起作用。该配方——将参考锚定在期望成员并细化——将RLHF风格RL的KL锚重新解释为选择旋钮,而不仅仅是稳定性约束。

英文摘要

Regularized self-play -- the family behind DeepNash's Stratego play -- drives a two-player zero-sum policy to a Nash equilibrium by best-responding to a slowly moving, entropy-regularized reference policy $ρ$. When the game has a polytope of value-equivalent equilibria, the regularizer silently breaks the tie: with a uniform reference it selects the maximum-entropy member, the I-projection of $ρ$ onto the Nash set. Can the reference be used to choose the equilibrium on purpose? On five exactly solvable games plus a 2-D polytope, with exact best responses and equivalence tests over independent seeds, anchoring the reference at a target member and refining steers self-play to that member with mean coordinate error 0.007 at median exploitability $5\times10^{-5}$, TOST-equivalent to the request within $\pm0.05$; the anchoring persists through refinement and follows the reference, not the initialization. Selection follows the reach-weighted I-projection (slope 0.969 [0.950, 0.987]). We report with equal emphasis where the story breaks: fixed off-manifold references cost 0.08-0.25 exploitability; stiff or flat families require a smaller mirror step, set by a pre-registered rule; boundary targets undershoot; curvature predicts where boundary saturation bites (rank correlation 0.90, p=0.037) while interior precision is curvature-independent. Table and MLP steering maps are equivalent within $\pm0.03$ at every target (30 seeds); matched control arms show attention's robust signature is excess seed variance, any systematic shift bounded at 0.018 and not significant. Against a best response the selection-robustness trade-off is degenerate: steering matters only against fixed, non-equilibrium opponents. The recipe -- anchor the reference at the desired member and refine -- reinterprets the KL anchor of RLHF-style RL as a selection knob, not only a stability leash.

Comments17 pages, 8 figures, 4 tables. Companion to arXiv:2606.28308 and arXiv:2607.17543. Fully reproducible: a single self-contained notebook regenerates every number, table, and figure

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑