CircuitLens:推理电路作为可验证奖励强化学习的数据选择信号
CircuitLens: Reasoning Circuits as Data Selection Signals for Reinforcement Learning with Verifiable Rewards
- School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院)
- Beijing Institute for General Artificial Intelligence(北京通用人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出电路推理得分(CRS),利用对比消融识别的注意力头作为数据选择信号,在RLVR中低参与度数据反而带来更好训练效果,表明数据选择具有情境依赖性。
AI中文摘要:
可验证奖励强化学习(RLVR)对模型训练所使用的问题集非常敏感,然而现有的选择标准——难度过滤、人工筛选、奖励轨迹评分——都将数据价值视为问题的内在属性,与将要学习该数据的模型无关。我们引入了电路推理得分(CRS),这是一种通过对比消融识别出的46个推理敏感注意力头导出的选择信号,在冻结的基础模型上进行单次前向传播即可计算,无需奖励标签或回滚。CRS与一个直观假设相悖,即更强的推理电路参与会产生更好的训练数据:在Qwen2.5-Math-7B上,参与度最低的十分位数在三个中等难度基准(GSM8K +2.0个百分点,OlympiadBench +1.6个百分点,Minerva +2.9个百分点)上优于随机选择,而参与度最高的十分位数增益较小,且与中间十分位数无法区分。该优势存在边界条件:在领域精选的数据池上,没有任何选择方法能与其他方法区分开来;在1.5B规模下,有效方向有所不同;且最低奖励的训练条件产生了最强的下游泛化能力。在所测试的Qwen2.5-Math设置中,RLVR数据选择似乎依赖于特定情境,而非简化为问题质量的静态排序。
英文摘要:
Reinforcement learning with verifiable rewards (RLVR) is sensitive to which problems a model trains on, yet existing selection criteria--difficulty filtering, hand-curation, reward-trajectory scoring--assess data value as an intrinsic property of problems, independent of the model that will learn from them. We introduce Circuit Reasoning Score (CRS), a selection signal derived from 46 reasoning-sensitive attention heads identified via contrastive ablation, computed in a single forward pass on the frozen base model without reward labels or rollouts. CRS runs against the intuitive hypothesis that stronger reasoning-circuit engagement produces better training data: on Qwen2.5-Math-7B, the lowest-engagement decile improves over random selection on three medium-difficulty benchmarks (GSM8K +2.0 pp, OlympiadBench +1.6 pp, Minerva +2.9 pp), while the highest-engagement decile gains less and is indistinguishable from the middle decile. The advantage has boundary conditions: on a domain-curated pool no selection method separates from the others; at 1.5B scale the useful direction differs; and the lowest-reward training condition produces the strongest downstream generalization. Within the Qwen2.5-Math settings tested, RLVR data selection appears regime-dependent rather than reducible to a static ranking of problem quality.