arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

偏好不确定性下的鲁棒纳什对齐

Robust Nash Alignment under Preference Uncertainty

Shihab Ahmed, Debamita Ghosh, David Tang, Yudan Wang, Alvaro Velasquez, Yue Wang

arXiv 2610.00715首次发表:更新:

发表机构

University of Central Florida; Arizona State University; University of Colorado Boulder(中佛罗里达大学; 亚利桑那州立大学; 科罗拉多大学博尔德分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对偏好不确定下的对齐脆弱性,提出鲁棒纳什对齐博弈框架,通过代理博弈与单循环算法实现最坏情况胜率下界保证,实验验证了收敛性与性能提升。

AI 中文摘要

基于偏好的对齐方法通常针对单一偏好模型进行优化,因此当成对偏好不确定时(如存在噪声、异质性或部署后发生偏移),这些方法可能变得脆弱。为解决这些问题,我们提出了鲁棒纳什对齐(Robust Nash Alignment),一个用于对齐不确定成对偏好的博弈论框架。我们的公式设定中,一个主学习器寻求一种策略,该策略在面对对抗性竞争对手以及位于名义偏好周围模糊集内的任何偏好核时,都能获得较大的最坏情况胜率。当模糊集捕捉了偏好的不确定性时,该博弈的鲁棒目标直接为最坏情况性能提供了经过认证的下界。然而,我们注意到该问题在计算上难以优化,为此,我们引入了一个涉及领导者策略、跟随者策略、对抗性核和对偶变量的四玩家原始-对偶代理博弈,并为其开发了一种单循环乐观镜像下降-上升算法。我们证明了该代理总是下界于截断的硬约束目标,量化了代理与硬目标之间的差距,并刻画了一个精确性条件,在该条件下代理能够恢复鲁棒目标。随后,我们证明了代理博弈对偶间隙的\\(\mathcal{O}(1/\sqrt{T})\\)平均迭代收敛性,这意味着对于原始鲁棒目标而言,该代理能产生一个接近最优的鲁棒策略。在受控表格博弈和具有不确定偏好的大语言模型对齐上的实验进一步验证了收敛理论,并显示出相对于名义基线的改进性能。

英文摘要

Preference-based alignment methods typically optimize against a single preference model, and can therefore be brittle when pairwise preferences are uncertain: noisy, heterogeneous, or shift after deployment. To address these issues, we propose Robust Nash Alignment, a game-theoretic framework for alignment to uncertain pairwise preferences. Our formulation has a major learner seeking a policy with a large worst-case win rate against both an adversarial competitor and any preference kernel lying in an ambiguity set around a nominal preference. When the ambiguity set captures the uncertainty in preferences, the resulting robust objective of the game directly yields a certified lower bound on worst-case performance. However, we note this problem is computationally challenging to optimize, and to address this, we introduce a four-player primal-dual proxy game involving the leader policy, follower policy, adversarial kernel, and dual variable, and develop a single-loop optimistic mirror descent-ascent algorithm for it. We show that the proxy always lower-bounds the truncated hard-constrained objective, quantify the proxy-to-hard gap, and characterize an exactness condition under which the proxy recovers the robust objective. We then prove an \(\mathcal{O}(1/\sqrt{T})\) average-iteration convergence for the proxy-game duality gap, which implies a near-optimal robust policy for the original robust objective. Experiments on controlled tabular games and LLM alignment with uncertain preference further validate the convergence theory and show improved performance over nominal baselines.

Comments38 pages, accepted at 2026 40th Advances in Neural Information Processing System (NeurIPS)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑