发表机构
UT Dallas; UT Austin; UC Santa Barbara(德克萨斯大学达拉斯分校; 德克萨斯大学奥斯汀分校; 加利福尼亚大学圣巴巴拉分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多指令遵循RL中偏向简单指令的探索偏差,本文提出两个衡量指标与含行为引导、稀缺感知奖励的两阶段框架,在三个基准上显著优于基线。
AI 中文摘要
强化学习(RL)已成为增强大型语言模型(LLM)指令遵循能力的强大范式。尽管现有训练方法取得了显著提升,但研究发现,当训练数据的一个提示中包含多个指令时,这些方法存在偏向简单指令的探索偏差。该偏差由两个主要原因导致:1)策略模型满足困难指令的初始能力过低,无法在RL训练中触发成功的探索,因此优化偏向简单指令;2)标准RL训练方法通常采用累积奖励(即满足的指令数量),将所有指令同等对待,这使得策略模型偏向于满足简单指令以获得相同的奖励。为解决这些问题,本文首先提出两个指标来衡量指令遵循中的探索偏差,随后引入一个两阶段框架来缓解该偏差:1)行为引导(Behavioral Bootstrapping),即RL训练前的轻量级拒绝采样微调阶段,用于激活困难指令;2)稀缺感知奖励(Scarcity-Aware Rewards),一种新的RL奖励函数,根据指令的经验稀缺性为其分配奖励。实验表明,所提出的指标与模型性能高度相关,且本文方法释放了RL训练的潜力:在三个可验证的指令遵循基准上,本文最优模型的性能显著优于基线模型。本文代码已发布在该httpsURL。
英文摘要
RL has emerged as a powerful paradigm for enhancing the instruction following capabilities of LLMs. While existing training recipes achieve substantial gains, we find that they suffer from exploration bias towards easy instructions when the training data has multiple instructions in a prompt. This bias is caused by two main reasons: 1) the policy model's initial ability to satisfy hard instructions is too low to trigger successful exploration during RL training, so the optimization is biased towards easy instructions; and 2) canonical RL training recipes typically employ a cumulative reward (the number of instructions fulfilled), treating all instructions equally, which biases the policy model towards fulfilling easy instructions to obtain the same amount of reward. To address these issues, we first propose two metrics to measure the exploration bias in instruction following and then introduce a two-stage framework to alleviate it: 1) Behavioral Bootstrapping, a lightweight rejection sampling fine-tuning stage before RL to activate hard instructions; and 2) Scarcity-Aware Rewards, a new RL reward function that assigns rewards to instructions based on their empirical scarcity. Experiments show that the proposed metrics are highly correlated with model performance, and our methods unleash the potential of RL training: our best models outperform the baselines by a significant margin across three verifiable instruction following benchmarks. We release codes at https://github.com/mianzhang/MulIF.
CommentsEMNLP 2026 Acceptance