arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

可检测边界处的后训练:一种用于微调的博弈论方法

Post-Training at the Edge of Detectability: A Game-Theoretic Approach to Fine-Tuning

Keegan Harris, Brian W. Lee, Ian Waudby-Smith, Philip Amortila, Nika Haghtalab, Michael I. Jordan

arXiv 2607.26358首次发表:更新:

发表机构

University of California, Berkeley; Inria; École Normale Supérieure(加州大学伯克利分校; 法国国家信息与自动化研究所; 巴黎高等师范学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出博弈论框架解决RL微调中KL正则化系数设置问题,将均衡系数归约为KL正则化RL目标,在Qwen3-8B等模型实验中实现良好奖励-保留权衡,可用于审计开源模型API提供商。

AI 中文摘要

强化学习(RL)微调广泛用于语言模型训练,以提升模型在目标任务上的性能,同时限制其与参考策略的偏差。平衡该权衡的标准方式是采用KL正则化的RL目标,但该公式本身并未提供设置正则化系数的原则性方法。实际中,该系数通常通过启发式选择或超参数搜索确定,这可能导致训练成本不必要的开销,或出现不合意的奖励-保留权衡。我们转而提出一种博弈论框架,为该权衡提供明确的统计解释。具体而言,我们研究一个序贯博弈:智能体选择策略以最大化累积奖励,而监控器随时间观察策略输出并测试其与参考策略的偏差。尽管并非源于同一视角,我们证明所得的均衡策略仍可表示为KL正则化RL问题的解,其中最优正则化参数可视为最大化单位统计区分度的奖励。利用凹-凸分式规划的经典结果,我们提供一种原则性方法,通过将其归约为KL正则化RL目标来学习该均衡系数,从而可灵活集成到标准微调流程中。在使用Qwen3-8B和Llama-3.2-1B的实验中,我们证明我们的方法在持续学习场景中可实现具有竞争力的奖励-保留权衡,并说明该框架可用于审计提供开源模型的API提供商。

英文摘要

Reinforcement learning (RL) fine-tuning is widely used in language model training to improve performance on a target task while limiting drift from a reference policy. A standard way to balance this trade-off is via a KL-regularized RL objective, although this formulation does not by itself provide a principled way to set the regularization coefficient. In practice, the coefficient is typically chosen heuristically or via hyperparameter search, which can lead to unnecessary overhead in training cost or undesirable reward-retention trade-offs. We instead propose a game-theoretic framework that gives this trade-off an explicit statistical interpretation. Specifically, we study a sequential game in which an agent chooses a policy to maximize cumulative reward while a monitor observes policy outputs over time and tests for deviations from the reference policy. Although not originating from the same perspective, we show that the resulting equilibrium policy can nonetheless be expressed as the solution to a KL-regularized RL problem for an optimal regularization parameter that can be viewed as maximizing reward per unit of statistical distinguishability. Drawing on classical results from concave-convex fractional programming, we provide a principled method for learning this equilibrium coefficient via reduction to the KL-regularized RL objective, thus allowing for flexible integration into standard fine-tuning pipelines. In experiments with Qwen3-8B and Llama-3.2-1B, we show that our methods result in competitive reward-retention trade-offs in a continual learning setting, and illustrate how our framework may be used to audit API providers serving open-source models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑