发表机构
Fudan University; Xingchen AGI Lab, China Telecom Artificial Intelligence Technology (Beijing) Co., Ltd.; Zhejiang University; Shanghai Jiao Tong University(复旦大学; 星辰AGI实验室,中国电信人工智能科技(北京)有限公司; 浙江大学; 上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对开放式强化学习中响应长度增长损害token效率的问题,提出质量门控长度优势塑造(QGLAS),让质量决定强化方向、长度仅塑造幅度,在约30%压缩下保留98.4%–102.0%的质量增益,优于基线。
AI 中文摘要
强化学习(RL)不仅改变语言模型所说的内容,还改变它们说多少,常常以牺牲token效率为代价来增加响应长度。在开放式RL中控制这种长度增长尤其具有挑战性,因为(i)响应长度与质量相互纠缠,(ii)开放式任务缺乏自然的成功边界来决定何时应优先考虑效率,以及(iii)密集的、分级奖励通常产生组内质量差异较小,使得质量引发的优势对奖励层面的长度塑造特别敏感,这种塑造可能扰动其幅度甚至逆转其符号。因此,我们采用一种不对称原则:质量应决定强化的方向,而长度仅应塑造其幅度。我们通过质量门控长度优势塑造(QGLAS)实例化这一原则,该机制首先仅从质量奖励计算优势,然后仅向较短的具有正优势的响应添加有界奖励,保持所有其他优势不变。奖励强度进一步适应组内质量分离,使得当质量偏好的响应相似时简洁性更重要,而当其质量差异明显时简洁性次要。在不同的模型家族、开放式基准和奖励来源中,QGLAS始终比代表性基线实现更强的质量-长度权衡。在大约30%压缩率下,QGLAS保留了质量驱动RL相对于基础模型所实现的宏观平均质量增益的98.4%–102.0%,而基线在类似压缩率下仅保留68.3%–75.5%。
英文摘要
Reinforcement learning (RL) changes not only what language models say, but also how much they say, often increasing response length at the cost of token efficiency. Controlling this length growth is particularly challenging in open-ended RL because (i) response length is entangled with quality, (ii) open-ended tasks lack a natural success boundary for deciding when efficiency should be prioritized, and (iii) dense, graded rewards often yield small within-group quality margins, making quality-induced advantages especially sensitive to reward-level length shaping, which can perturb their magnitudes and even reverse their signs. We therefore adopt an asymmetric principle: quality should determine the direction of reinforcement, while length should only shape its magnitude. We instantiate this principle with Quality-Gated Length Advantage Shaping (QGLAS), which first computes advantages from quality rewards alone, then adds bounded bonuses only to shorter positive-advantage responses, leaving all other advantages unchanged. The bonus strength is further adapted to within-group quality separation, allowing conciseness to matter more when quality-favored responses are similar and less when their quality differences are clear. Across different model families, open-ended benchmarks, and reward sources, QGLAS consistently achieves a stronger quality--length trade-off than representative baselines. At approximately 30% compression, QGLAS retains 98.4--102.0% of the macro-average quality gains achieved by quality-only RL over the base model, compared with 68.3--75.5% for these baselines at comparable compression.
Comments22 pages. Preprint, under review