弹性地平线:在智能体强化学习中探索有效交互边界
Elastic Horizon: Discovering the Effective Interaction Frontier in Agentic Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
针对智能体强化学习中交互地平线扩展的收益递减问题,提出弹性地平线闭环控制器,基于成功轨迹长度第90百分位数动态调整地平线,在多个基准上实现最优性能并节省高达25%的令牌。
中文摘要 AI 辅助
扩展交互地平线(即每个回合中环境交互的最大次数)能够提升大型语言模型智能体在长时程任务上的表现,而基于课程学习、逐步扩展地平线的方法优于固定地平线的替代方案。然而,现有的调度方案是开环的:它们单调地增加地平线,直到手动指定的最大值,且缺乏检测进一步扩展何时不再有帮助的机制。我们提出了有效交互边界假说:一个动态边界,超过该边界后,额外的交互带来的收益递减,而成本却线性增长。随后,我们引入了弹性地平线(Elastic Horizon),一种闭环控制器,通过成功轨迹长度的第90百分位数来追踪该边界。在AppWorld和BFCL上,固定地平线的扫描显示出明显的饱和平台;弹性地平线从欠容量和过容量两种初始化状态出发,均能将地平线稳定在饱和带内,在7B和14B骨干模型上取得了最佳成功率,并节省了高达25%的每步轨迹令牌。我们的工作将范式从如何扩展交互地平线转变为何时停止扩展。
英文摘要
Scaling the interaction horizon-the maximum number of environment interactions per episode-improves LLM agents on long-horizon tasks, and curriculum-based methods that progressively expand the horizon outperform fixed-horizon alternatives. However, existing schedules are open-loop: they monotonically increase the horizon until a manually specified maximum, with no mechanism to detect when further expansion stops helping. We propose the effective interaction frontier hypothesis: a dynamic boundary beyond which additional interactions yield diminishing returns while cost grows linearly. We then introduce Elastic Horizon, a closed-loop controller that tracks this boundary via the 90th percentile of successful trajectory lengths. On AppWorld and BFCL, fixed-horizon sweeps reveal clear saturation plateaus; Elastic Horizon stabilizes the horizon inside the saturation band from both under- and over-capacity initializations, attains the best success rates across 7B and 14B backbones, and saves up to 25% of per-step trajectory tokens. Our work shifts the paradigm from how to scale interaction horizons to when to stop scaling.
发表机构
- Qwen Business Unit of Alibaba(阿里巴巴通义千问业务部)
- University of Science and Technology of China(中国科学技术大学)
机构由 AI 辅助整理,请以论文原文为准。