超越不确定性:用于自演进推理课程的多求解器分歧奖励
Beyond Uncertainty: Multi-Solver Disagreement Rewards for Self-Evolving Reasoning Curricula
浏览论文内容
中文总结 AI 辅助
该研究针对自演进推理框架中单求解器奖励的瓶颈,提出多求解器分歧奖励方法,使求解器在竞赛数学基准上平均提升1.34个点,为自博弈推理系统课程生成提供互补可扩展信号。
中文摘要 AI 辅助
自演进推理框架会训练挑战者(Challenger)生成暴露求解器(Solver)弱点的问题,从而在无需人工数据的情况下创建自适应课程。然而,现有方法采用单个求解器的采样不确定性作为挑战者的奖励,这会产生一个根本性瓶颈:当求解器对挑战者的问题分布变得自信时,所有采样答案会完全一致,导致奖励归零,使挑战者缺乏学习信号。关键的是,这种单模型奖励无法区分真正简单的问题和仅与单个求解器学习偏差对齐的问题。我们提出一种多求解器分歧奖励,使用在模型容量和采样温度上存在差异的异构集成模型。对集成模型针对每个问题的多数答案计算归一化香农熵,明确奖励求解器产生冲突解决方案的问题——将难度捕获为模型间的分歧而非模型内的采样方差。这种更丰富的梯度使挑战者能够发现针对真实能力边界的问题,生成的课程迫使下游求解器开发能跨问题类型泛化的稳健推理策略。我们的方法是可直接替换的奖励函数,无需修改框架或额外数据。对Qwen3-4B的实验显示,在以分歧挑战者问题训练的求解器上,竞赛数学基准(MATH-500、AMC、奥林匹克)的平均性能提升了1.34个点,表明多求解器分歧为自博弈推理系统的课程生成提供了互补且可扩展的信号。
英文摘要
Self-evolving reasoning frameworks train a Challenger to generate questions exposing a Solver's weaknesses, creating adaptive curricula without human data. However, existing approaches use a single solver's sampling uncertainty as the Challenger's reward. This creates a fundamental bottleneck: as the solver grows confident on the Challenger's question distribution, all sampled answers converge identically, collapsing the reward to zero and starving the Challenger of learning signal. Critically, this single-model reward cannot distinguish genuinely easy questions from those that merely align with one solver's learned biases. We propose a multi-solver disagreement reward using a heterogeneous ensemble varying in model capacity and sampling temperature. A normalized Shannon entropy over the ensemble's per-question plurality answers explicitly rewards questions where solvers produce conflicting solutions---capturing difficulty as inter-model divergence rather than intra-model sampling variance. This richer gradient enables the Challenger to discover questions targeting true capability boundaries, producing a curriculum that forces downstream Solvers to develop robust reasoning strategies generalizing across problem types. Our approach is a drop-in reward function replacement requiring no framework modifications or additional data. Experiments with Qwen3-4B show that Solvers trained on disagreement-Challenger questions achieve +1.34 points average improvement on competition-math benchmarks (MATH-500, AMC, Olympiad), suggesting that multi-solver disagreement provides a complementary and scalable signal for curriculum generation in self-play reasoning systems.