发表机构
Shanghai Jiao Tong University(上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对LLM智能体长期搜索任务中展开预算分配不合理问题,提出基于信息增益的展开策略优化(IGRPO),通过按节点信息性分配预算进行树结构展开,并利用诱导教师分布指导策略优化,实验验证该方法在相同预算下优于基线。
AI 中文摘要
强化学习已成为改进大型语言模型(LLM)智能体在长期搜索任务中的一种有前景的范式,智能体在获得最终结果前需做出一系列中间决策。然而,现有方法存在关键局限:展开预算分配时未明确评估中间状态的效用,导致大量计算花费在低价值状态。本文提出基于信息增益的展开策略优化(IGRPO),将中间状态信息性作为展开收集的组织原则。IGRPO通过根据节点级信息性分配扩展预算进行预算感知树结构展开,更频繁扩展信息丰富分支,抑制无前景分支。还证明基于信息增益的展开在轨迹上诱导出明确的极限教师分布,产生清晰的策略优化目标,统一了自适应树结构探索与有原则的策略学习。在七个具有挑战性的搜索增强问答基准测试上的实验表明,IGRPO在相同展开预算约束下持续优于强大基线,验证了利用诱导教师分布指导长期搜索智能体策略优化的有效性。
英文摘要
Reinforcement learning (RL) improves large language model (LLM) agents on long-horizon search tasks that require multiple intermediate decisions before a final outcome. However, rollout budgets are often allocated without assessing intermediate-state utility, which can waste computation on unpromising branches. We propose Information Gain-based Rollout Policy Optimization (IGRPO), a framework that organizes rollout collection around intermediate-state informativeness. Specifically, IGRPO performs budget-aware tree-structured rollouts in which expansion probabilities depend on node-level informativeness, allowing informative branches to receive more computation while less informative branches are expanded less frequently within a fixed rollout budget. By directly shaping how training trajectories are generated, IGRPO induces a limiting teacher distribution over search trajectories that favors higher cumulative informativeness. The resulting distribution provides an explicit policy optimization target, connecting adaptive rollout collection with principled policy learning. Experiments on seven search-augmented question answering benchmarks show that IGRPO achieves higher average accuracy than strong baselines on both 3B and 7B backbones under comparable rollout budgets, supporting informativeness-guided trajectory generation for training search agents. Code is available at https://github.com/e3trange/IGRPO.