Dr. Free:自进化搜索智能体无需难度奖励
Dr. Free: You Don't Need Difficulty Rewards for Self-Evolving Search Agents
浏览论文内容
中文总结 AI 辅助
Dr. Free提出首个无需难度奖励的自进化搜索框架,通过知识图谱关系链生成多跳问题并优化证据必要性,减少训练时间7倍以上,在七个开放域问答基准上超越现有方法。
中文摘要 AI 辅助
当前用于训练搜索智能体的无数据自进化方法的一个核心局限是依赖基于难度的提议者奖励。这些方法通过奖励提议者生成挑战协同进化求解器的问题,将求解器难度作为问题质量的代理指标。然而,仅凭难度不足以区分需要跨段落证据的问题与可通过更简单捷径回答的问题。此外,测量难度需要对每个候选问题重复运行求解器,导致巨大的计算成本。在本文中,我们引入了\methodname,这是首个消除基于难度的提议者奖励并直接针对捷径上下文优化证据必要性的自进化搜索框架。Dr. Free从知识图谱中采样关系链,并将其与对齐的段落配对,为问题生成赋予显式的多跳结构。仅当完整证据段落下的目标答案似然超过所有评估的捷径上下文下的最大似然时,生成的问题才获得正信息增益奖励。由于该信号基于教师强制似然计算,它无需通过率估计,并将提议者训练时间减少超过7倍。在七个开放域问答基准上的实验表明,Dr. Free优于先前的无数据搜索智能体和监督基线,在多跳问答基准上取得了大幅改进。
英文摘要
A central limitation of current data-free self-evolution methods for training search agents is their reliance on difficulty-based proposer rewards. These methods reward a proposer for generating questions that challenge a co-evolving solver, using solver difficulty as a proxy for question quality. Yet difficulty alone is insufficient to distinguish questions that require cross-passage evidence from those that are answerable via simpler shortcuts. In addition, measuring difficulty demands repeated solver rollouts for every candidate question, leading to substantial computational costs. In this paper, we introduce \methodname, the first self-evolving search framework that eliminates difficulty-based proposer rewards and directly optimizes for evidence necessity relative to shortcut contexts. Dr. Free samples relational chains from a knowledge graph and pairs them with aligned passages, giving question generation an explicit multi-hop structure. A generated question receives a positive information-gain reward only when the likelihood of the target answer under the complete evidence passages exceeds the maximum likelihood under all evaluated shortcut contexts. Because this signal is computed from teacher-forced likelihoods, it removes the need for pass-rate estimation and reduces proposer training time by over $7\times$. Experiments on seven open-domain QA benchmarks show that Dr. Free outperforms prior data-free search agents and the supervised baseline, with large improvements on multi-hop QA benchmarks.
发表机构
- VCIP, School of Computer Science, Nankai University(南开大学计算机学院VCIP)
- Kuaishou Technology(快手科技)
- Xiamen University(厦门大学)
机构由 AI 辅助整理,请以论文原文为准。