发表机构
Santa Clara University; Rochester Institute of Technology(圣克拉拉大学; 罗切斯特理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
HindSearch是面向GRPO的轨迹级事后批判方法,通过冻结评判器利用标准答案为失败轨迹提供搜索动作的在线蒸馏信号,在7基准测试上使Qwen2.5-3B-Instruct的平均EM达39.4%,优于现有基线。
AI 中文摘要
搜索增强型大语言模型智能体通常采用二元精确匹配奖励训练,会丢弃失败轨迹中关于失败原因的大部分信息。我们提出HindSearch,一种面向GRPO的事后自蒸馏流程:每次rollout后,冻结的评判器会利用标准答案对每条失败轨迹撰写简短批判,该批判为学生智能体的搜索动作提供辅助的在线蒸馏信号。在使用Qwen2.5-3B-Instruct的标准7基准测试套件上,HindSearch达到39.4%的平均精确匹配率(EM),优于现有搜索-RL基线方法。若移除评判器对标准答案的访问权限,性能提升会大部分消失,这表明事后机制是性能改善的来源。
英文摘要
Search-augmented LM agents are typically trained with a binary exact-match reward, which throws away most of what a failed trajectory tells us about why it failed. We introduce HindSearch, a hindsight self-distillation procedure for GRPO: after each rollout, a frozen judge writes a short critique of every failed trajectory using the gold answer, and the critique supplies an auxiliary on-policy distillation signal on the student's search actions. On the standard seven-benchmark suite with Qwen2.5-3B-Instruct, HindSearch reaches 39.4% average EM, outperforming prior search-RL baselines. Removing the judge's access to the gold answer erases most of the gain, isolating hindsight as the source of the improvement.