arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向搜索智能体的上下文信息策略优化

Contextual Information Policy Optimization for Search Agents

Xingyu Guo, Wei Chen, Linlin Yang, Baochang Zhang

arXiv 2608.06128首次发表:更新:

AI 中文总结

针对现有搜索智能体方法存在的先验驱动推理问题,提出CIPO框架,通过回合级信用分配对齐策略与证据使用,无需额外标注或奖励模型,在7个基准上性能优异。

AI 中文摘要

搜索智能体通过让大型语言模型在多步推理过程中获取和使用外部证据,将其扩展至静态参数记忆之外。对于涉及复杂或动态信息的知识密集型任务,其可靠性不仅取决于检索相关证据,还取决于利用证据指导后续推理。然而,现有方法主要奖励最终答案正确性或中间进展,未直接评估检索后的行动是否基于检索到的证据。这种不一致会催生先验驱动的推理:智能体基于内部知识形成结论,仅将检索用于确认结论,导致确认偏差和低效推理。为此,我们提出Contextual Information Policy Optimization(CIPO,上下文信息策略优化),这是一种面向证据的强化学习框架,可明确将策略优化与外部证据使用对齐。CIPO为受检索信息影响的推理行动分配密集的回合级信用,同时将该证据使用信号与全局结果奖励结合以实现稳定学习。CIPO阻止脱离证据的猜测,促进检索事实可指导或修正后续推理的轨迹。重要的是,CIPO既不需要人类过程标注,也不需要额外的奖励模型。在7个域内和域基准上的大量实验表明,CIPO降低了先验驱动推理的流行度,并在大多数任务上取得了优异性能。

英文摘要

Search agents extend large language models beyond static parametric memory by enabling them to acquire and use external evidence during multi-step reasoning. For knowledge-intensive tasks involving complex or evolving information, their reliability depends not only on retrieving relevant evidence but also on using it to guide subsequent reasoning. However, existing methods primarily reward final-answer correctness or intermediate progress, without directly assessing whether post-retrieval actions are grounded in the retrieved evidence. This misalignment encourages prior-driven reasoning: agents form conclusions based on internal knowledge and use retrieval mainly to confirm them, resulting in confirmation bias and inefficient evidence use. To address this issue, we propose Contextual Information Policy Optimization (CIPO), an evidence-oriented reinforcement learning framework that explicitly aligns policy optimization with external evidence use. CIPO assigns dense, turn-level credit to reasoning actions influenced by retrieved information, while combining this evidence-use signal with a global outcome reward to preserve answer correctness. With this manner, CIPO discourages evidence-detached guesses and promotes reasoning trajectories in which retrieved facts can guide or revise subsequent reasoning. Importantly, CIPO requires neither human process annotations nor an additional reward model. Extensive experiments on seven in-domain and out-of-domain benchmarks show that CIPO reduces the prevalence of prior-driven reasoning and achieves excellent performance on most tasks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑