按你所说采样:让语言模型对齐到它们所陈述的分布
Sample What You Say: Aligning Language Models to Sample the Distributions They State
- New York University(纽约大学)
- Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对语言模型能陈述分布却无法从中采样的问题,提出基于最大均值差异的见证优势作为GRPO奖励,显著降低采样分布与目标的偏差并保持通用能力。
AI中文摘要:
语言模型越来越多地被用于从指定分布中采样,例如模拟调查受访者或生成合成数据。经过指令微调的模型能够正确陈述这样的分布,但仍然无法从中采样。提示和改变解码方式只能部分减少这种不匹配,这促使我们使用策略优化进行训练。组相对策略优化(GRPO)是解决这一问题的自然选择,因为它已经为每个提示采样一组轨迹,并且可以将该组的经验分布与目标分布进行比较。然而,将整个组作为一个整体进行评分会给每个轨迹相同的奖励。组相对中心化随后将所有优势设为零,模型便无法获得学习信号。为了给每个轨迹提供其自身的信号,我们引入了见证优势(witness advantage),这是一种基于最大均值差异(MMD)的逐轨迹优势。它训练模型去匹配有限结果集上的目标分布。模型分布与目标分布之间的MMD具有一个见证函数,该函数衡量每个结果是过度产生还是产生不足。每个轨迹的优势估计其结果的负见证值,因此,如果某个结果在组中产生不足,该轨迹会获得奖励;如果某个结果在组中过度产生,该轨迹会受到惩罚。见证优势根据组的结果计数以闭式形式计算,我们将其用作GRPO中的奖励。在未见过的目标分布上,使用见证优势进行训练显著降低了与目标分布的总变差距离,同时基本保持了模型的通用能力。
英文摘要:
Language models are increasingly used to sample from a specified distribution, for instance, to simulate survey respondents or generate synthetic data. Instruction-tuned models can state such a distribution correctly and still fail to sample from it. Prompting and changes to decoding reduce this mismatch only partly, which motivates training with policy optimization. Group relative policy optimization (GRPO) is a natural fit for this problem because it already samples a group of rollouts per prompt, and the group's empirical distribution can be compared with the target. However, scoring the group as a whole gives every rollout the same reward. Group-relative centering then sets all advantages to zero, and the model receives no learning signal. To give each rollout its own signal, we introduce the witness advantage, a per-rollout advantage derived from maximum mean discrepancy (MMD). It trains a model to match a target distribution over a finite set of outcomes. The MMD between the model's distribution and the target has a witness function that measures how over- or under-produced each outcome is. Each rollout's advantage estimates the negative witness at its outcome, so a rollout is rewarded for an outcome the group under-produces and penalized for one it over-produces. The witness advantage is computed in closed form from the group's outcome counts, and we use it as the reward in GRPO. On unseen target distributions, training with the witness advantage substantially reduces the total variation distance to the target while largely preserving the model's general capabilities.