arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

推理时纳什对齐

Inference-Time Nash Alignment

Hadi Hosseini, Debmalya Mandal, Duohan Zhang

arXiv 2609.08082首次发表:更新:

发表机构

Penn State University; University of Warwick(宾夕法尼亚州立大学; 华威大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对偏好微调成本高且需参数访问的问题,提出推理时纳什对齐方法,通过二人零和博弈求纳什均衡,两种算法达理论下界且实验优于基础策略。

AI 中文摘要

基于偏好的微调方法,如RLHF和DPO,需要大量的计算资源和庞大的偏好数据集。它们还需要直接访问模型参数,而许多最先进的模型并不提供这些参数。推理时对齐提供了一种无需更新模型参数的成本效益高的替代方案。然而,现有的推理时方法依赖于在Bradley-Terry假设下推导出的标量奖励模型,该模型无法表示一般偏好。继最近关于广义偏好微调的研究之后,在这项工作中,我们开创了对一般偏好下推理时对齐的研究。我们将该问题表述为获取两个策略之间的二人零和博弈的纳什均衡。我们提出了两种算法:Best-of-Nash(BoN)和Nash Mirror Descent(NMD)。我们证明了这两种算法都能达到与问题下界匹配的对偶间隙。在实证方面,我们在三个数据集上实现了这两种方法,结果表明我们的方法显著优于基础策略,并收敛到微调模型的性能。此外,我们的结果显示NMD在正则化参数变化时保持稳健。

英文摘要

Preference-based fine-tuning methods such as RLHF and DPO require substantial compute and large preference datasets. They also need direct access to the model parameters which are not provided by many state-of-the art models. Inference-time alignment offers a cost-effective alternative without updating model parameters. However, existing inference-time methods rely on a scalar reward model derived under a Bradley-Terry assumption, which cannot represent general preferences. Following recent work on fine-tuning with generalized preferences, in this work, we initiate the study of inference-time alignment under general preferences. We formulate the problem as obtaining a Nash equilibrium of a two-player zero-sum game between policies. We propose two algorithms: Best-of-Nash (BoN) and Nash Mirror Descent (NMD). We prove that both algorithms achieve a duality gap that matches the problem lower bound. Empirically, we implement the two methods on three datasets, which shows that our methods substantially outperform the base policy, converging to the performance of the fine-tuned models. Moreover, our results show that NMD remains robust across the regularization parameter.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑