arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

单一策略,任意预算:通过强化学习内化预算感知搜索

One Policy, Any Budget: Internalizing Budget-Aware Search via Reinforcement Learning

Xiaowei Sun, Jin Li, Yili Hong, Yikun Fu, Yanghua Xiao

arXiv 2609.00813首次发表:更新:

发表机构

College of Computer Science and Artificial Intelligence, Fudan University; College of Software Engineering, Southeast University; School of Artificial Intelligence, Shanghai Jiao Tong University(复旦大学计算机与人工智能学院; 东南大学软件学院; 上海交通大学人工智能学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对现有LLM搜索智能体无法适配变化预算约束的问题,提出AnySearch框架,通过两阶段课程强化学习训练,在7个问答基准上实现了跨预算规模的性能提升与泛化能力。

AI 中文摘要

尽管强化学习已使基于大语言模型(LLM)的搜索智能体能够调用外部工具,但现有方法在固定预算下训练,无法在部署时约束条件变化时进行适配。我们提出AnySearch,一个通过训练框架和课程强化学习实现单一策略在任意预算约束下执行预算感知搜索的框架。第一阶段,我们训练智能体时注入显式预算状态,并使用结构化推理提示,指导其在线性衰减的预算下进行高效分配。第二阶段,移除训练框架,智能体在自适应采样的预算约束下自主运行,匹配推理条件。两个阶段均采用复合奖励优化,该奖励通过绝对和相对信号将答案准确率与预算效率结合,其中自适应权重会放大高准确率查询的效率信号,衰减低准确率查询的效率信号。在7个通用和多跳问答基准上的大量实验表明,我们的方法在所有预算规模下均优于基线,可泛化到训练范围之外的未见约束,且在无过多令牌开销的情况下实现了更优的工具生产力。我们的代码可在此https URL获取。

英文摘要

While reinforcement learning has enabled LLM-based search agents to invoke external tools, existing methods train under fixed budgets and cannot adapt when constraints vary at deployment. We propose AnySearch, a framework that enables a single policy to perform budget-aware search under any budget constraint through a training scaffold and curriculum reinforcement learning. In the first phase, we train the agent with explicit budget state injection and structured reasoning prompts that guide efficient allocation under linearly decaying budgets. In the second phase, the scaffold is removed and the agent learns to operate autonomously under adaptively sampled budget constraints, matching inference conditions. Both phases are optimized with a composite reward that couples answer accuracy with budget efficiency through absolute and relative signals, where an adaptive weight amplifies the efficiency signal for high-accuracy queries and attenuates it for low-accuracy ones. Extensive experiments on seven general and multi-hop QA benchmarks show that our method outperforms baselines across all budget scales, generalizes to unseen constraints beyond the training range, and achieves superior tool productivity without excessive token overhead. Our code is available at https://github.com/xwsun01/AnySearch.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑