发表机构
Zhejiang University; ZJU-UIUC Institute, Zhejiang University; Tencent Inc.(浙江大学; 浙江大学ZJU-UIUC学院; 腾讯公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对搜索智能体RL微调的可靠性问题,提出CAS框架,通过自适应检索与策略加权结合CP技术,提升推理准确率并减少冗余搜索,构建高可靠高效智能体范式。
AI 中文摘要
搜索智能体在强化学习(RL)微调过程中面临严重的可靠性危机。启发式Top-K检索常导致关键证据丢失或噪声混入,而渐进式RL引发的过度自信会导致答案幻觉与冗余搜索。为构建高可靠智能体,我们引入保形预测(CP)并提出保形智能体搜索(CAS)框架,该框架在检索与训练两端均建立可靠性保障:在检索端,自适应预测集(APS,一种特定的CP实现)将统计覆盖率转化为动态文档截断,以构建大小自适应的预测集;在训练端,自适应保形推理(ACI,一种动态CP算法)构建具有可控覆盖率的预测集,以量化答案置信度,随后将其用于在组相对策略优化(GRPO)目标内惩罚低置信度轨迹,确保模型仅从可靠轨迹中学习。在单跳与多跳问答数据集上的实验表明,该框架显著提升推理准确率,同时大幅减少冗余工具调用,建立了高可靠且高效的智能体范式。我们的代码可在此URL获取。
英文摘要
Search Agents face a severe reliability crisis during reinforcement learning (RL) fine-tuning. Heuristic Top-K retrieval often causes critical evidence loss or noise inclusion, while over-confidence induced by progressive RL leads to hallucinated answers and redundant searches. To build highly reliable agents, we introduce Conformal Prediction (CP) and propose Conformalized Agentic Search (CAS). This framework establishes reliability guarantees on both the retrieval and training sides: on the retrieval side, an Adaptive Prediction Set (APS), a specific CP realization, translates statistical coverage into dynamic document truncation to construct prediction sets that are adaptive in size; on the training side, Adaptive Conformal Inference (ACI), a dynamic CP algorithm, dynamically constructs prediction sets with controllable coverage to quantify answer confidence, which is then used to penalize low-confidence trajectories within the Group Relative Policy Optimization (GRPO) objective, ensuring the model learns only from reliable ones. Experiments across single-hop and multi-hop QA datasets demonstrate that our framework significantly improves reasoning accuracy while drastically reducing redundant tool invocations, establishing a highly reliable and efficient agent paradigm. Our code is available at https://github.com/S1llyBird/CAS.
Comments21 pages, including figures, tables, and appendix