arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13866cs.LG

过滤指标是安全关键的:形状奖励下组相对强化学习中的幻影优势

The Filter Metric is Safety-Critical: Phantom Advantages in Group-Relative RL under Shaped Rewards

  • Fudan University(复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

Juntao Yu

AI总结:

本研究揭示组相对强化学习中过滤指标与谓词不匹配导致幻影优势,提出应使用任务结果信号进行过滤,实验证明可显著提升模型性能。

AI中文摘要:

组相对策略优化(GRPO及其衍生方法)可以通过动态采样丢弃无对比的 rollout 组,而实际实现中暴露了一个可配置的过滤指标。我们识别并量化了复合形状奖励下的指标-谓词不匹配问题。当过滤依据形状化训练分数而非任务结果时,全失败组保留组内非零方差并通过谓词;随后标准差归一化将失败组中的形状差异提升为全尺寸的幻影优势。在受控的 GSM8K 对比实验(Qwen2.5-1.5B,LoRA)中,无过滤和形状分数过滤分别以 EM 0.080 ± 0.112 和 0.040 ± 0.008 结束,而二元结果过滤在每臂四次运行(三次默认种子重跑和一次种子123运行;均值 ± 样本标准差)中保持 0.754 ± 0.005。在 verl 的原生 recipe/dapo 训练器上,保持模型、数据、奖励和训练器固定且仅更改指标,分数臂在观察到的 40 步中无需批次补充,最终 EM 为 0.160;准确率臂在 40 步中补充了 29 次,最终为 0.763。两者均使用相同的自定义形状奖励钩子和未修改的训练器/过滤代码。先前工作建立了形状诱导的放大和全失败过滤;我们的贡献在于隔离了指标-谓词语义不匹配,并直接检测了原生删除/补充遥测。在测试的正系数 lambda ∈ {0.1, 0.3, 0.5} 下,不安全臂崩溃;探索性单次运行单元在 MATH 上 1.5B/7B 规模及 GSPO 下重现了失败,而禁用标准差归一化则避免了观察到的崩溃。复合奖励下的过滤应使用语义独立于形状化的任务结果信号。

英文摘要:

Group-relative policy optimization (GRPO and descendants) can discard no-contrast rollout groups through dynamic sampling, while practical implementations expose a configurable filter metric. We identify and quantify a metric-predicate mismatch under composite shaped rewards. When filtering follows the shaped training score rather than the task outcome, all-fail groups retain nonzero within-group spread and pass the predicate; standard-deviation normalization then promotes shaping differences among failures to full-size phantom advantages. In a controlled GSM8K comparison (Qwen2.5-1.5B, LoRA), no filtering and shaped-score filtering end at EM 0.080 +/- 0.112 and 0.040 +/- 0.008, whereas binary-outcome filtering holds 0.754 +/- 0.005 across four runs per arm (three default-seed reruns and one seed-123 run; mean +/- sample SD). On verl's native recipe/dapo trainer, holding model, data, reward and trainer fixed and changing only the metric, the score arm requires no batch refill in any of 40 observed steps and ends at EM 0.160; the accuracy arm refills in 29/40 steps and ends at 0.763. Both use the same custom shaped-reward hook and unmodified trainer/filter code. Prior work established shaping-induced amplification and all-fail filtering; our contribution isolates the metric-predicate semantic mismatch and directly instruments native deletion/refill telemetry. Across tested positive coefficients lambda in {0.1, 0.3, 0.5}, unsafe arms collapse; exploratory one-run cells reproduce the failure at 1.5B/7B on MATH and under GSPO, while disabling standard-deviation normalization avoids the observed collapse. Filtering under a composite reward should use a task-outcome signal whose semantics are independent of shaping.

补充信息

↑