发表机构
Moore Threads AI(摩尔线程人工智能)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM强化学习中响应偏离策略导致的概率比率符号抵消问题,提出取消感知响应掩码(CARM),通过取对数比率绝对值防止漂移被掩盖,并在数学推理与代码基准上显著提升性能。
AI 中文摘要
近年来,强化学习(RL)在大语言模型(LLM)的后训练中得到了快速应用,在数学推理和代码生成方面取得了显著进展。然而,在实际系统中,策略更新以及 rollout 与训练引擎之间的差异可能使采样响应偏离策略(off-policy)。序列级掩码通过决定整个响应是否应参与优化来解决这一不匹配问题。一种常见的掩码规则使用采样令牌概率比率的长度归一化几何均值。其带符号的对数比率可能在不同位置相互抵消,从而掩盖了显著的双向策略漂移。我们提出了取消感知响应掩码(CARM),这是一种序列级掩码,在平均之前先取每个令牌对数比率的绝对值,从而防止相反的概率变化相互抵消。我们证明了被接受的响应满足一个联合约束,该约束涉及采样令牌比率超出规定区间的比例以及它们超出边界时的平均对数距离。在数学推理和代码生成上的实验表明,与几何均值掩码相比,CARM 在 AIME 2024/2025/2026 和 BeyondAIME 上的平均 mean@16 最高提升了 3.13 个百分点,并且在四个代码基准上将平均 pass@1 相对于最强评估基线提高了 2.88 个百分点。这些发现支持 CARM 作为一种理论上合理且有效的方法,用于大语言模型强化学习中的响应级离策略控制。
英文摘要
Recent years have witnessed the rapid adoption of reinforcement learning (RL) in large language model (LLM) post-training, with substantial gains in mathematical reasoning and code generation. In practical systems, however, policy updates and differences between rollout and training engines can make sampled responses off-policy. Sequence-level masking addresses this mismatch by deciding whether an entire response should contribute to optimization. A common masking rule uses the length-normalized geometric mean of sampled token probability ratios. Its signed log-ratios can cancel across positions, concealing substantial bidirectional policy drift. We propose \emph{Cancellation-Aware Response Masking} (CARM), a sequence-level mask that takes the absolute value of each token log-ratio before averaging, preventing opposing probability changes from canceling. We prove that accepted responses satisfy a joint bound on the fraction of sampled-token ratios outside a prescribed band and their mean log-distance beyond its boundaries. Experiments on mathematical reasoning and code generation show that CARM improves mean@16 averaged over AIME 2024/2025/2026 and BeyondAIME by up to $3.13$ percentage points over geometric-mean masking, and increases average pass@1 across four code benchmarks by $2.88$ points over the strongest evaluated baseline. These findings support CARM as a theoretically grounded and effective method for response-level off-policy control in LLM reinforcement learning.
Comments28 pages, 11 figures, 5 tables