arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

语言模型对齐的渐近行为与记忆效应

The Asymptotics of Language Model Alignment with Memory

Haricharan Balasundaram, V. Arvind Rameshwar

arXiv 2610.01828首次发表:更新:

发表机构

Georgia Institute of Technology; IIT Madras(佐治亚理工学院; 印度理工学院马德拉斯分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文扩展了语言模型对齐的渐近接近性结果至马尔可夫输出序列,并完整刻画了有限长度(m=1)下两种对齐方法KL散度为零的条件。

AI 中文摘要

语言模型(LM)对齐广泛旨在将给定的LM $Q$ 扰动为对齐的LM $q$,使得 i) $q$ 和 $Q$ 产生的输出在概率上“接近”,ii) $q$ 具有比 $Q$ 更高的期望奖励。两种常见的LM对齐技术是:KL约束强化学习,它需要知道LM分布且计算成本高;以及best-of-$n$算法,它仅需从LM中采样。Yang等人的工作建立了两种对齐方法产生的分布之间的渐近接近性,针对LM输出的$m$长度独立同分布(i.i.d.)标记序列,在$m$趋于无穷大的极限下。然而,i.i.d.假设不能代表实际LM,其输出序列通常具有记忆性。在本文中,我们将渐近接近性结果扩展到LM输出的$m$长度标记序列为马尔可夫链的情形。此外,对于有限长度输出序列——特别是当$m=1$时——我们提供了LM分布和奖励函数的完整刻画,使得两种对齐方法产生的分布之间的KL散度为零——这是Yang等人首次提出的问题。

英文摘要

Language model (LM) alignment broadly aims to perturb a given LM $Q$ into an aligned LM $q$ such that i) the outputs produced by $q$ and $Q$ are 'close' in probability, ii) $q$ has a higher expected reward than $Q$. Two common techniques for LM alignment are: KL-constrained RL, which requires knowledge of the LM distribution and is computationally expensive, and the best-of-$n$ algorithm, which requires only sampling from the LM. The work of Yang et al. established asymptotic closeness between the distributions produced by the two alignment methods for an $m$--length i.i.d. token sequence output by the LM, in the limit as $m$ increases to infinity. However, the i.i.d. assumption is not representative of practical LMs, whose output sequences often have memory. In this paper, we extend the asymptotic closeness result to the case when the $m$--length token sequence outputted by the LM is Markovian. Further, for finite-length output sequences -- particularly, when $m=1$ -- we provide a complete characterization of LM distributions and reward functions for which the KL-divergence between the distributions produced by the two alignment methods is zero -- a question first posed in Yang et al.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑