AdaKP:面向推理强化学习的在线自适应知识点选择
AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
研究针对强化学习在竞赛级数学中奖励稀疏问题,提出AdaKP在线自适应知识点选择方法,通过熵代理评分,结合三种机制及验证门,改进标准训练器,在竞赛数学基准上优于静态选择基线。
中文摘要 AI 辅助
具有可验证奖励的强化学习是在大语言模型中引发推理的强大范式,但在竞赛级数学上存在严重的奖励稀疏性。常见补救方法是将原子知识点(KPs)注入提示中。现有方法要么离线固定选择,要么仅缩放注入文本的整体数量。我们引入AdaKP,一个在强化学习训练过程中重新选择每个问题的KP子集的在线选择器。其核心是一个熵代理,通过它引起的下一个token熵的减少来对KP进行评分。三种轻量级机制使该信号可在线使用。AdaKP还贡献了一个飞行前验证门。作为标准DAPO+GRPO训练器的完全加法分支,AdaKP在所有八个竞赛数学基准上优于强大的静态选择基线,成本可忽略不计,将在线、经过验证的KP子集选择定位为面向推理强化学习的实用且未充分探索的方向。
英文摘要
Reinforcement learning with verifiable rewards is a powerful paradigm for eliciting reasoning in large language models, yet it suffers from severe reward sparsity on competition-level mathematics. A common remedy injects atomic knowledge points (KPs) - short natural-language hints distilled from gold solutions - into the prompt. Existing methods, however, either fix this selection once offline or merely scale the monolithic quantity of injected text, leaving untouched the most informative axis of choice: which subset of atomic KPs to inject, and when. We introduce AdaKP, an online selector that re-chooses each problem's KP subset over the course of RL training. At its core is an entropy proxy that scores a KP by the reduction in next-token entropy it induces - a single inexpensive forward pass, with a provable bound on its truncation bias - in place of expensive rollout-based estimation. Three lightweight mechanisms make this signal usable online: a momentum smoother that absorbs per-step noise, a retirement-and-revival manager that prunes weak KPs while preserving exploration, and an adaptive scheduler that front-loads re-evaluations into early training. AdaKP further contributes a pre-flight validation gate that certifies the proxy against a leave-one-out ground truth before any expensive run is launched, turning method-level risk into a falsifiable check. Realized as a fully additive fork of a standard DAPO+GRPO trainer with no optimizer changes, AdaKP improves over a strong static-selection baseline on all eight competition-mathematics benchmarks at negligible added cost, positioning online, validated KP-subset selection as a practical and as-yet under-explored axis for reasoning-oriented reinforcement learning.