发表机构
Harbin Institute of Technology(哈尔滨工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究多轮证据阅读语言模型代理训练不稳定问题,提出CIGPO方法,通过方差注入策略,利用冻结参考模型对数似然边际增加作逐轮信号,避免奖励方差崩溃,在HotpotQA实验中取得更好效果。
AI 中文摘要
仅通过结果进行强化学习来训练多轮证据阅读代理是不稳定的,因为中间轮次获得的直接奖励很少。在Qwen2.5 - 3B - Instruct的HotpotQA实验中,GRPO最初有所改进(标准F1为0.430),但随后崩溃为100%格式违规输出。训练日志诊断揭示了零优势锁定机制。我们提出了一种方差注入策略,通过为中间证据阅读轮次分配逐轮奖励来防止奖励分布崩溃。上下文信息增益策略优化(CIGPO)使用冻结参考模型对真实答案的对数似然的边际增加作为逐轮信号来实现此策略。通过IG和F1奖励的单独归一化以及IG权重课程,CIGPO在3B规模的HotpotQA上达到了标准F1为0.518,而最佳GRPO检查点为0.430,最终GRPO检查点为0.000。CIGPO在整个训练过程中保持有意义的奖励方差并避免零优势锁定。这些结果表明奖励方差崩溃是仅结果GRPO的一种具体失败模式,并且轮级IG奖励可以在HotpotQA设置中防止这种情况。
英文摘要
Training multi-turn evidence-reading agents with outcome-only reinforcement learning is unstable because intermediate turns receive little direct credit. In HotpotQA experiments with Qwen2.5-3B-Instruct, GRPO initially improves (standard F1 0.430) but subsequently collapses to 100% format-violating outputs. Training-log diagnosis reveals a zero-advantage lock-in mechanism: all sampled trajectories receive the minimum format penalty (-2.0), group-relative advantages vanish, and the policy-gradient loss becomes zero--an optimization deadlock. We propose a variance-injection strategy: by assigning per-turn rewards to intermediate evidence-reading turns, we prevent the group reward distribution from collapsing to a single value--preserving the variation that GRPO's group-relative advantage requires. Contextual Information-Gain Policy Optimization (CIGPO) implements this strategy using the marginal increase in the frozen reference model's log-likelihood of the ground-truth answer as the per-turn signal. With separate normalization of IG and F1 rewards and an IG-weight curriculum, CIGPO reaches a standard F1 of 0.518 on HotpotQA at the 3B scale (from 0.252 base; +105%), compared with 0.430 for the best GRPO checkpoint and 0.000 for the final GRPO checkpoint. CIGPO maintains meaningful reward variance and avoids zero-advantage lock-in throughout training. These results identify reward-variance collapse as a concrete failure mode of outcome-only GRPO and show that turn-level IG rewards can prevent it in this HotpotQA setting.