arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CompassPlay:对提议者引导求解器的方向给予奖励

CompassPlay: Rewarding the Proposer for Where It Moves the Solver

Sophia Xiao Pu, Ximeng Sun, Jiang Liu, Jialian Wu, Emad Barsoum, Zicheng Liu, William Yang Wang

arXiv 2609.32228首次发表:更新:

发表机构

University of California, Santa Barbara; Advanced Micro Devices(加州大学圣巴巴拉分校; 超威半导体)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

CompassPlay通过梯度对齐奖励提议者,引导任务生成以提升训练效率,在编码和定理证明任务中显著提高性能并减少计算成本。

AI 中文摘要

在自博弈中,提议者生成可验证的任务来训练求解器。提议者的奖励通常取决于求解器的成功率,但同等困难的任务在训练价值上可能有所不同。我们引入了CompassPlay,一种通过梯度对齐来奖励提议者的自博弈方法。该奖励倾向于那些其求解器损失梯度与代表目标能力的参考任务梯度对齐的任务。它利用学习进展的一阶近似,并在无需额外训练求解器的情况下对每个合格任务进行评分。我们的实验展示了性能和训练效率的提升。在编码自博弈中,使用Qwen2.5-Coder-7B,一个小型外部参考集引导任务生成。CompassPlay在领域内编码上的平均准确率比AZR的难度奖励提高了1.5个百分点,在领域外数学上提高了2.7个百分点。在Lean4定理证明中,CompassPlay以少40%的GPU小时数达到了难度基线150次迭代的累积覆盖率。

英文摘要

In self-play, a proposer generates verifiable tasks to train a solver. Proposer rewards often depend on the solver's success rate, but equally difficult tasks can differ in their training value. We introduce CompassPlay, a self-play method that rewards the proposer through gradient alignment. The reward favors tasks whose solver loss gradients align with those of reference tasks representing the target capabilities. It draws on a first-order approximation to learning progress and scores each eligible task without additional solver training. Our experiments show gains in performance and training efficiency. In coding self-play with Qwen2.5-Coder-7B, a small external reference set guides task generation. CompassPlay improves average accuracy over AZR's difficulty reward by 1.5 percentage points on in-domain coding and 2.7 on out-of-domain mathematics. In Lean4 theorem proving, CompassPlay matches the difficulty baseline's 150-iteration cumulative coverage with 40\% fewer GPU-hours.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑