Search-G1:基于表征内在奖励的接地搜索智能体
Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards
浏览论文内容
中文总结 AI 辅助
该研究提出Search-G1框架,通过两个经干预校准的读数构成的表征内在奖励,改善了搜索增强语言智能体的接地性与搜索成本的权衡,在多基准和模型规模上验证了其有效性。
中文摘要 AI 辅助
结合搜索增强的语言智能体应仅在必要时检索外部信息,并基于检索到的证据生成答案。现有外部奖励要么提供稀疏的结果监督,要么提供来自过程标注和大语言模型(LLM)评判的更丰富反馈。结果奖励易于扩展,但无法区分接地检索与冗余搜索;而更丰富的信号则需要昂贵的标注或训练期间的推理。基于策略侧信号(如熵、似然或信息增益)的内部奖励是分级的,评估成本低廉,但主要反映模型置信度而非证据接地。我们提出Search-G1,这是一个基于表征的内在奖励框架,通过两个经干预校准的读数来衡量智能体答案的操作接地性:提示-状态读数预测闭卷知识的充分性,其补集定义策略相关的检索必要性;答案-承诺读数通过答案阶段对证据删除的敏感性来估计对证据的依赖。两者结合,当估计需要检索且答案对证据敏感时,为正确的搜索轨迹提供额外奖励;当闭卷知识足够时,奖励正确的直接答案;并惩罚重复搜索。校准后,奖励评分在策略优化期间既不需要过程标注,也不需要LLM作为评判的推理。由于强化学习会改变策略表征,Search-G1会定期基于最新检查点的轨迹重新拟合两个读数,使奖励与策略共同进化。在多个基于搜索的问答基准和两种模型规模上的实验表明,Search-G1改善了接地性-搜索成本的权衡,在具有竞争力的任务准确率下产生了更短的响应侧轨迹。代码可在this https URL获取。
英文摘要
Search-augmented language agents should retrieve external information only when necessary and ground their answers in retrieved evidence. Existing external rewards provide either sparse outcome supervision or richer feedback from process annotations and LLM judges. Outcome rewards scale readily but cannot distinguish grounded retrieval from redundant search, whereas richer signals require costly annotation or inference during training. Internal rewards based on policy-side signals such as entropy, likelihood, or information gain are graded and inexpensive to evaluate, yet mainly reflect model confidence rather than evidence grounding. We propose Search-G1, a representation-based intrinsic reward framework that measures the operational grounding of an agent's answers through two intervention-calibrated readouts. A prompt-state readout predicts closed-book sufficiency, whose complement defines policy-relative retrieval necessity; an answer-commit readout estimates evidence reliance from answer-stage sensitivity to evidence deletion. Together, they provide additional credit to correct searched trajectories when retrieval is estimated necessary and the answer is evidence-sensitive, favor correct direct answers when closed-book knowledge suffices, and penalize repeated search. After calibration, reward scoring requires neither process annotations nor LLM-as-judge inference during policy optimization. Because reinforcement learning changes policy representations, Search-G1 periodically refits both readouts on trajectories from the latest checkpoint, allowing the reward to co-evolve with the policy. Experiments across multiple search-based question-answering benchmarks and two model scales show that Search-G1 improves the grounding--search-cost trade-off, producing shorter response-side trajectories at competitive task accuracy. Code is available at https://github.com/Rosy0912/Search-G1.
发表机构
- Fudan University(复旦大学)
- Tencent(腾讯)
- Nanjing University(南京大学)
- Nanyang Technological University(南洋理工大学)
- Shanghai Jiao Tong University(上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。