arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

QLPO:用于长度感知策略优化的象限加权采样

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization

Siwei Chen, Siqi Chen, Xupeng Miao, Bin Cui

arXiv 2607.21793首次发表:更新:

发表机构

School of Computer Science, Peking University; Beijing Key Laboratory of Software and Hardware Cooperative Artificial Intelligence Systems, Peking University; Department of Electronic Engineering, Tsinghua University; Institute of Computational Social Science, Peking University (Qingdao)(北京大学计算机科学学院; 北京大学软硬件协同人工智能系统北京市重点实验室; 清华大学电子工程系; 北京大学(青岛)计算社会科学研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对大型推理模型长思维链响应致高延迟成本问题,提出QLPO方法,通过先过度生成候选响应,再按特定规则重采样训练组,实现隐式长度控制,在多模型中改善准确率与长度权衡,减少响应长度并保持推理性能。

AI 中文摘要

近期大型推理模型在强化学习中常生成长思维链响应,导致推理延迟和部署成本高。现有响应长度控制方法依赖显式长度惩罚或额外控制模块,需精细调整且可能影响推理质量。我们提出用于长度感知策略优化的象限加权采样(QLPO),这是GRPO基于重采样的简单变体,在不修改奖励函数的情况下引入隐式长度控制。QLPO先过度生成候选响应,然后通过保留经验正确/错误率,同时偏向短正确响应和长错误响应来对训练组进行重采样。这重塑了训练分布并隐式鼓励更短的模型输出。在参数从15亿到320亿的模型中,包括基础模型和强推理模型,QLPO持续改善了准确率与长度的权衡。它在保持推理性能的同时将响应长度减少了30%至70%。这些结果表明结构化重采样为高效推理提供了有效且稳健的方法。

英文摘要

Recent large reasoning models often develop long chain-of-thought responses during reinforcement learning (RL), resulting in high inference latency and deployment cost. Existing methods for response length control typically rely on explicit length penalties or additional control modules, which require careful tuning and may compromise reasoning quality. We propose Quadrant-weighted Sampling for Length-aware Policy Optimization (QLPO), a simple resampling-based variant of GRPO that introduces implicit length control without modifying the reward function. QLPO first over-generates candidate responses and then resamples the training group by preserving the empirical correct/incorrect ratio while favoring short correct responses and long incorrect responses. This reshapes the training distribution and implicitly encourages shorter model outputs. Across models ranging from 1.5B to 32B parameters, including both base models and strong reasoning models, QLPO consistently improves the accuracy-length trade-off. It reduces response length by 30% to 70% while preserving reasoning performance. These results suggest that structured resampling provides an effective and robust approach to efficient reasoning.

Comments18 pages, 7 figures, 6 tables. Accepted at COLM 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑