EviSD:面向搜索增强智能体的证据条件自蒸馏
EviSD: Evidence-Conditioned Self-Distillation for Search-Augmented Agents
浏览论文内容
中文总结 AI 辅助
该研究针对搜索增强智能体的轨迹级信用问题,提出EviSD证据条件自蒸馏框架,在7个问答基准及3种骨干模型上,其宏观平均精确匹配优于最强对比方法1.3-2.3个百分点。
中文摘要 AI 辅助
基于结果的强化学习使搜索增强型语言智能体能够从可验证的最终答案中学习,但其轨迹级信用无法区分多轮搜索过程中各个动作的贡献。我们提出EviSD,这是一种证据条件自蒸馏框架,将实例级支持证据作为搜索动作的特权信息,将黄金答案作为答案动作的补充特权信息。训练期间,学生从原始上下文采样动作,而同一模型在动作对齐的上下文中作为特权教师对这些动作重新评分。EviSD将分离的师生差距转化为对结果衍生的GRPO优势的有界校正,且仅应用于生成的动作跨度。该设计在保留结果奖励确定的更新方向的同时,将特权指导本地化,无需辅助蒸馏目标,也无需在推理时进行任何更改。在涵盖模型规模和生成的7个问答基准及3种骨干模型上,EviSD在所有评估设置中实现了最高的宏观平均精确匹配,比最强的对比方法高出1.3至2.3个百分点,同时仅调整6.7%至15.1%的响应标记。代码可在https URL获取。
英文摘要
Outcome-based reinforcement learning enables search-augmented language agents to learn from verifiable final answers, but its trajectory-level credit cannot distinguish the contributions of individual actions in a multi-turn search process. We propose EviSD, an evidence-conditioned self-distillation framework that uses instance-level supporting evidence as privileged information for search actions and golden answers as complementary privilege for answer actions. During training, the student samples actions from the original context, while the same model re-scores them as a privileged teacher under an action-aligned context. EviSD converts the detached teacher--student gap into a bounded correction to the outcome-derived GRPO advantage and applies it only to generated action spans. This design localizes privileged guidance while preserving the update direction determined by the outcome reward, without an auxiliary distillation objective or any change at inference time. Across seven question-answering benchmarks and three backbones spanning model scales and generations, EviSD achieves the highest macro-average Exact Match in all evaluated settings, outperforming the strongest compared methods by 1.3--2.3 points while modulating only 6.7%--15.1% of response tokens. Code is available at https://github.com/JiananXie/EviSD.
发表机构
- Ant Group(蚂蚁集团)
- NLPR, MAIS, CASIA(中国科学院自动化研究所模式识别国家重点实验室、多模态人工智能系统实验室)
机构由 AI 辅助整理,请以论文原文为准。