发表机构
The Hong Kong University of Science Technology; Yuanbao Team, Tencent(香港科技大学; 腾讯元宝团队)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出 HATCH 框架,通过在线加权退火提示轨迹贡献和梯度投影解决提示奖励偏移,实现无辅助的自改进推理,在数学基准上显著超越现有方法。
AI 中文摘要
组相对策略优化(GRPO)通过比较每个查询的多个解决方案 rollout 的验证奖励来改进语言模型的推理能力。然而,困难的训练查询可能只产生错误的 rollout,使得 GRPO 缺乏奖励对比或学习信号。先前的基于提示的方法从解决方案证据中构建辅助提示,并用它们重新解决失败的查询,从而恢复学习信号。然而,尽管这些轨迹是在评估时不可用的辅助条件下生成的,它们通常被视为普通解决方案轨迹。我们发现了提示奖励偏移:恢复的奖励对比可以将策略更新集中在提示轨迹上,从而限制了在没有提示的情况下的改进。这也产生了一个权衡:增加提示轨迹可以加速早期学习,但会在后期加剧奖励偏移。为了解决这个问题,我们提出了 HATCH(提示退火自教学),一个在线单策略框架,通过生成和使用自己的提示来学习,从而在没有辅助的情况下改进推理。为了缓解提示奖励偏移,我们引入了在线加权来退火提示轨迹的贡献。然而,学习生成提示可能与改进查询解决相冲突。因此,我们使用梯度投影来移除提示生成更新中的对立成分。这些设计共同支持自我改进,使策略能够为自己创造学习机会,并将其转化为没有提示的更强推理。我们在数学推理基准上评估了我们的方法,并在 Llama-3.2-1B-Instruct 上比最先进的方法高出 1.02 个百分点,在 Qwen3-1.7B 上高出 2.84 个百分点,在 Qwen3-8B 上高出 4.32 个百分点。
英文摘要
Group Relative Policy Optimization (GRPO) improves language-model reasoning by comparing verified rewards among multiple solution rollouts for each query. However, difficult training queries can yield only incorrect rollouts, leaving GRPO with no reward contrast or learning signal. Prior hint-based methods construct auxiliary hints from solution evidence and use them to re-solve failed queries, recovering learning signal. Yet the resulting trajectories are typically treated as ordinary solution trajectories despite being generated under an assisted condition unavailable at evaluation. We discover hinted reward shift: recovered reward contrast can concentrate policy updates on hinted trajectories, limiting improvement without hints. This also creates a trade-off: increasing hinted trajectories can accelerate early learning but intensify reward shift later. To address this problem, we propose HATCH (Hint-Annealed Self-Teaching), an online single-policy framework that learns from both generating and using its own hints to improve reasoning without assistance. To mitigate hinted reward shift, we introduce online weighting to anneal the contribution of hinted trajectories. However, learning to generate hints can conflict with improving query solving. We therefore use gradient projection to remove the opposing component of hint-generation updates. Together, these designs support self-improvement by enabling the policy to create learning opportunities for itself and turn them into stronger reasoning without hints. We evaluate our method on mathematical reasoning benchmarks and outperform state-of-the-art methods by 1.02 pp on Llama-3.2-1B-Instruct, 2.84 pp on Qwen3-1.7B, and 4.32 pp on Qwen3-8B.