AI 中文总结
本研究针对搜索智能体检索过程中的等价崩溃问题,提出图结构检索框架Harness-G,结合结构化非近视信用机制,在六个问答基准上显著优于现有方法。
AI 中文摘要
强化学习(RL)搜索智能体通常将检索建模为自由形式的自然语言查询生成,并通过最终答案奖励优化多轮交互。当前研究主要通过更密集或更结构化的信用信号改进训练,但很少探究在策略-环境接口处的检索是否得到了恰当的建模。我们在Search-R1训练过程中观察到明显的检索混叠现象:针对同一问题的不同rollout会持续生成不同的查询字符串,但其累积的证据集合却越来越重叠。我们将这一现象称为检索等价崩溃;在此状态下,轨迹在检索决策方面趋近于效用等价,使得组内回报几乎没有有效的检索区分度。为解决该问题,我们提出Harness-G,这是一种重新设计了上述接口的图结构检索框架。它将自由形式的查询生成重新建模为有限动作选择:策略选择一个证据句子或实体,或选择回答;而环境则构建菜单、跟踪检索状态,并验证和执行每个选择。该接口减少了语言混叠,使相同状态下的备选方案可直接比较。基于此接口,我们引入结构化非近视信用(Structured Non-myopic Credit, SNC),它使用冻结的答案评分器来比较所选动作与其备选方案,并将下游收益分配给产生这些收益的早期动作。在六个问答基准上,Harness-G在两种评估的模型规模下均取得了最高的平均F1值,在1.5B参数规模下比最强基线Graph-R1高出10.74个百分点,在3B参数规模下高出3.98个百分点。
英文摘要
Reinforcement learning (RL) search agents commonly model retrieval as free-form natural-language query generation and optimize multi-turn interactions using final-answer rewards. Current studies mainly improve training with denser or more structured credit signals, but rarely examine whether retrieval is properly formulated at the policy-environment interface. We observe pronounced retrieval aliasing during Search-R1 training: rollouts for the same question continue to generate distinct query strings, yet their accumulated evidence sets increasingly overlap. We call this phenomenon retrieval-equivalence collapse; in this regime, trajectories approach utility equivalence with respect to retrieval decisions, leaving within-group returns with little effective retrieval contrast. To address this problem, we propose Harness-G, a graph-structured retrieval framework that redesigns this interface. It reformulates free-form query generation as finite action selection: the policy selects an evidence sentence or entity, or chooses to answer, while the environment constructs the menu, tracks retrieval state, and validates and executes each choice. This interface reduces linguistic aliasing and makes same-state alternatives directly comparable. Building on this interface, we introduce Structured Non-myopic Credit (SNC), which uses a frozen answer scorer to compare the selected action with its alternatives and assigns downstream gains to the earlier actions that enabled them. Across six QA benchmarks, Harness-G achieves the highest average F1 at both evaluated model scales, outperforming the strongest baseline, Graph-R1, by 10.74 points at 1.5B and 3.98 points at 3B.
CommentsCode:https://github.com/7HHHHH/Harness-G