IGSD:用于搜索智能体的环境验证后见自蒸馏
IGSD: Environment-Verified Hindsight Self-Distillation for Search Agents
浏览论文内容
中文总结 AI 辅助
IGSD通过环境反馈验证查询令牌提议,利用配对信息增益作为软权重进行自蒸馏,在七个问答基准上显著提升搜索智能体的精确匹配准确率。
中文摘要 AI 辅助
在线策略自蒸馏无需外部教师即可强化智能体训练:基于特权后见的策略为其自身的非特权轨迹提供逐步指导。然而,对于搜索智能体而言,后见可能使教师偏好于从学生状态出发无法改善检索的查询。现有方法要么直接蒸馏这种偏好,要么使用模型内部分数进行过滤,但两种策略均未验证查询实际执行的检索结果。我们提出信息增益门控自蒸馏(IGSD),该方法在蒸馏前通过环境反馈验证在线策略的令牌提议。将每个查询令牌视为微动作,IGSD将教师的令牌提议与学生的采样令牌补全为匹配的查询,并使用相同的检索器从相同的失败状态执行两者。共享的反事实控制解释了查询条件导致的答案可能性偏移,因此它们的差异——执行后的配对信息增益——为检索到的文档提供了相对效用对比。IGSD将此对比作为候选对蒸馏的正向软权重,同时保持GRPO目标不变并将验证限制在训练阶段。在七个单跳和多跳问答基准上,IGSD在3B和7B策略下分别达到42.8%和47.0%的宏平均精确匹配准确率,且无需推理时验证。这些结果支持环境验证的后见作为搜索智能体可靠动作级监督的有效方法。
英文摘要
On-policy self-distillation densifies agent training without external teachers: a policy conditioned on privileged hindsight provides step-level guidance for its own unprivileged rollouts. For search agents, however, hindsight can make the teacher prefer a query that does not improve retrieval from the student's state. Existing methods either distill this preference directly or filter it with model-internal scores, but neither strategy verifies the query's executed retrieval consequence. We propose Information-Gain-Gated Self-Distillation (IGSD), which verifies on-policy token proposals with environment feedback before distilling them. Treating each query token as a micro-action, IGSD completes the teacher's token proposal and the student's sampled token into matched queries and executes both from the same failed state with the same retriever. Shared counterfactual controls account for query-conditioned shifts in answer likelihood, so their difference, the executed paired information gain, provides a relative utility contrast for the retrieved documents. IGSD uses this contrast as a positive-only soft weight for candidate-pair distillation, while leaving the GRPO objective unchanged and confining verification to training. Across seven single-hop and multi-hop QA benchmarks, IGSD reaches macro-average exact-match accuracies of 42.8% and 47.0% with 3B and 7B policies, respectively, without inference-time verification. These results support environment-verified hindsight as an effective approach to reliable action-level supervision for search agents.
发表机构
- University of Science and Technology of China(中国科学技术大学)
- Tsinghua University(清华大学)
- Meituan(美团)
机构由 AI 辅助整理,请以论文原文为准。