arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24794cs.AI

CAFE:自改进搜索智能体需要协同演化的反馈

CAFE: Self-Improving Search Agents Need Co-Evolving Feedback

  • Fudan University(复旦大学)
  • Tencent(腾讯)

机构由 AI 辅助整理,请以论文原文为准。

Boyang Liu, Senjie Jin, Peixin Wang, Zhangyue Yin, Yibo Wang, Yuhao Zhou, Zhihao Zhang, Xinbing Liang, Shizheng Zhu, Yuhui Wang, Jingqi Tong, Dingwei Zhu, Zhihe… 展开作者

Boyang Liu, Senjie Jin, Peixin Wang, Zhangyue Yin, Yibo Wang, Yuhao Zhou, Zhihao Zhang, Xinbing Liang, Shizheng Zhu, Yuhui Wang, Jingqi Tong, Dingwei Zhu, Zhiheng Xi, Jiazheng Zhang, Clive Bai, Clarenceai, Blaze Chen, Tao Gui, Qi Zhang, Xuanjing Huang

AI总结:

该研究提出 CAFE 框架,通过共享参数模型耦合搜索智能体与评判者的协同演化反馈,在七个及六个域外搜索基准上均优于现有 RL 智能体,减少答案幻觉且性能持续提升。

AI中文摘要:

结果监督的搜索智能体学习何时及如何检索证据,但终端奖励既无法定位中间错误,也无法在这些错误累积前重定向当前轨迹。将纠正反馈视为学习到的轨迹内干预措施,可将两种角色耦合:智能体必须决定何时请求和使用反馈,而评判者必须从结果混淆的 rollout 中推断有用的纠正,这些 rollout 的失败模式会随着智能体的改进而变化。我们引入 CAFE(Coupled Agent--Feedback Evolution,耦合智能体-反馈演化),这是一个共享参数模型在搜索智能体和评判者角色间交替的框架。CAFE 从围绕基础智能体自身失败构建的轨迹中初始化反馈条件恢复,随后耦合在线和离线优化。在线强化学习(RL)期间,比较反馈估计使用提示级别的调用-跳过成功差距来塑造请求回报,而感知反馈的优势塑形会重新加权反馈前后的 token 优势。离线阶段,源自 rollout 的偏好优化从匹配的成功和失败轨迹中学习反馈。在七个智能体搜索基准上,CAFE 平均优于所评估的基于 RL 的搜索智能体,在全部六个域外基准上保持其增益,并减少答案层面的幻觉。单侧 ablation 显示,仅改进智能体或仅改进评判者最终会达到性能平台,而交替两者的更新则会持续提升性能。这些发现表明,自改进搜索智能体需要与其所指导的策略协同演化的反馈。

英文摘要:

Reliable search requires more than acquiring external evidence. An agent must also recognize and recover from errors as its trajectory unfolds. In-trajectory feedback provides a mechanism for such recovery by diagnosing where the search has drifted and redirecting subsequent reasoning steps. This is particularly important in long-horizon search, where an early directional error may receive no immediate corrective signal and can compound across later steps. Making such feedback learnable, however, creates a coupled problem: the agent must learn when to request and use feedback, while the critic must learn corrections from outcome-confounded rollouts as the agent's failure patterns evolve. We introduce CAFE (Coupled Agent--Feedback Evolution), a framework in which a shared-parameter model alternates between search-agent and critic roles. CAFE initializes feedback-conditioned recovery from trajectories built around the base agent's own failures, then couples online and offline optimization. During online RL, a comparative feedback estimate uses a prompt-level call--skip success gap to shape request returns, while feedback-aware advantage shaping reweights token advantages before and after feedback. Offline, rollout-derived preference optimization learns feedback from matched successful and unsuccessful trajectories. On seven agentic search benchmarks, CAFE outperforms the evaluated RL-based search agents on average, retains its gains across all six out-of-domain benchmarks, and reduces answer-level hallucinations. One-sided ablations show that improving only the agent or only the critic eventually plateaus, whereas alternating the two updates continues to improve performance. These findings suggest that a self-improving search agent needs feedback that co-evolves with the policy it guides.

↑