CAFE:自改进搜索智能体需要协同演化的反馈
CAFE: Self-Improving Search Agents Need Co-Evolving Feedback
- Fudan University(复旦大学)
- Tencent(腾讯)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究提出 CAFE 框架,通过共享参数模型耦合搜索智能体与评判者的协同演化反馈,在七个及六个域外搜索基准上均优于现有 RL 智能体,减少答案幻觉且性能持续提升。
AI中文摘要:
结果监督的搜索智能体学习何时及如何检索证据,但终端奖励既无法定位中间错误,也无法在这些错误累积前重定向当前轨迹。将纠正反馈视为学习到的轨迹内干预措施,可将两种角色耦合:智能体必须决定何时请求和使用反馈,而评判者必须从结果混淆的 rollout 中推断有用的纠正,这些 rollout 的失败模式会随着智能体的改进而变化。我们引入 CAFE(Coupled Agent--Feedback Evolution,耦合智能体-反馈演化),这是一个共享参数模型在搜索智能体和评判者角色间交替的框架。CAFE 从围绕基础智能体自身失败构建的轨迹中初始化反馈条件恢复,随后耦合在线和离线优化。在线强化学习(RL)期间,比较反馈估计使用提示级别的调用-跳过成功差距来塑造请求回报,而感知反馈的优势塑形会重新加权反馈前后的 token 优势。离线阶段,源自 rollout 的偏好优化从匹配的成功和失败轨迹中学习反馈。在七个智能体搜索基准上,CAFE 平均优于所评估的基于 RL 的搜索智能体,在全部六个域外基准上保持其增益,并减少答案层面的幻觉。单侧 ablation 显示,仅改进智能体或仅改进评判者最终会达到性能平台,而交替两者的更新则会持续提升性能。这些发现表明,自改进搜索智能体需要与其所指导的策略协同演化的反馈。
英文摘要:
Reliable search requires more than acquiring external evidence. An agent must also recognize and recover from errors as its trajectory unfolds. In-trajectory feedback provides a mechanism for such recovery by diagnosing where the search has drifted and redirecting subsequent reasoning steps. This is particularly important in long-horizon search, where an early directional error may receive no immediate corrective signal and can compound across later steps. Making such feedback learnable, however, creates a coupled problem: the agent must learn when to request and use feedback, while the critic must learn corrections from outcome-confounded rollouts as the agent's failure patterns evolve. We introduce CAFE (Coupled Agent--Feedback Evolution), a framework in which a shared-parameter model alternates between search-agent and critic roles. CAFE initializes feedback-conditioned recovery from trajectories built around the base agent's own failures, then couples online and offline optimization. During online RL, a comparative feedback estimate uses a prompt-level call--skip success gap to shape request returns, while feedback-aware advantage shaping reweights token advantages before and after feedback. Offline, rollout-derived preference optimization learns feedback from matched successful and unsuccessful trajectories. On seven agentic search benchmarks, CAFE outperforms the evaluated RL-based search agents on average, retains its gains across all six out-of-domain benchmarks, and reduces answer-level hallucinations. One-sided ablations show that improving only the agent or only the critic eventually plateaus, whereas alternating the two updates continues to improve performance. These findings suggest that a self-improving search agent needs feedback that co-evolves with the policy it guides.