arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ActiveSaddler:用于智能体框架优化的自动课程学习

ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization

Sungho Park, Wonjoong Kim, Jue Zhang, Wook-Shin Han, Pengfei Gao, Chanyoung Park, Yongqiang Yao, Rao Fu, Elsie Nallipogu, Qingwei Lin, Victor Rühle

arXiv 2610.00906首次发表:更新:

发表机构

POSTECH; KAIST; Microsoft(浦项科技大学; 韩国科学技术院; 微软)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ActiveSaddler提出将框架优化中的训练课程建模为非平稳老虎机,动态构建失败模式目标并自适应平衡探索与利用,在GAIA2和Terminal-Bench 2.0上显著提升测试Pass@1。

AI 中文摘要

自动框架优化可以通过迭代更新提示、工具接口和控制逻辑来显著改进大语言模型智能体。然而,现有方法主要优化框架的更新方式,而很大程度上固定了生成驱动这些更新的反馈的训练场景。随着框架的演进,对进一步优化最有用的场景可能会发生变化,这表明训练课程本身应随框架一同调整。我们将框架优化的这一缺失维度表述为一个自动课程学习问题,并引入ActiveSaddler。ActiveSaddler将演进的课程建模为一个具有动态实例化优化目标的非平稳老虎机。它将重复出现的失败抽象为可复用的失败模式臂,估计进一步针对每种模式进行优化的潜在学习进展,并自适应地平衡重新审视已知弱点与探索未见场景以发现新弱点。优化结果持续更新已发现失败模式的集合及其优先级,使课程与框架共同演进。在GAIA2和Terminal-Bench 2.0上的实验表明,ActiveSaddler持续发现更强的框架,在测试Pass@1上分别比使用优化前固定场景顺序的相同框架优化器提高了4.4和7.5个百分点。消融实验进一步表明,这些收益依赖于动态构建优化目标、估计其演进效用以及平衡持续优化与新失败发现。这些结果共同确立了自动课程学习作为框架优化的一个新的关键优化维度。

英文摘要

Automated harness optimization can substantially improve LLM agents by iteratively updating their prompts, tool interfaces, and control logic from execution feedback. However, existing methods primarily optimize how the harness is updated while largely fixing which training scenarios generate the feedback that drives those updates. As the harness evolves, the scenarios most useful for further optimization can change, suggesting that the training curriculum itself should adapt alongside the harness. We formulate this missing dimension of harness optimization as an automated curriculum learning problem and introduce ActiveSaddler. ActiveSaddler models the evolving curriculum as a non-stationary bandit with dynamically instantiated optimization targets. It abstracts recurring failures into reusable failure-pattern arms, estimates the potential learning progress from further targeting each pattern, and adaptively balances revisiting known weaknesses with exploring unseen scenarios for new ones. Optimization outcomes continually update both the set of discovered failure patterns and their priorities, allowing the curriculum to co-evolve with the harness. Experiments on GAIA2 and Terminal-Bench 2.0 show that ActiveSaddler consistently discovers stronger harnesses, improving test Pass@1 by 4.4 and 7.5 percentage points over the same harness optimizer using a scenario order fixed before optimization, respectively. Ablations further show that these gains depend on dynamically constructing optimization targets, estimating their evolving utility, and balancing continued optimization with new failure discovery. Together, these results establish automated curriculum learning as a new crucial optimization dimension for harness optimization.

Comments37 pages, 16 figures. Project website and code: https://autosaddler-projectpage.github.io/activesaddler/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑