AutoSaddler:基于智能体执行轨迹的持久化更新的自动工具优化
AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces
浏览论文内容
中文总结 AI 辅助
AutoSaddler是基于智能体执行轨迹的自动工具优化框架,通过离线学习结合故障诊断等技术,在多个基准上提升智能体性能9.0-10.0个百分点,为构建更可靠的智能体系统提供了新方向。
中文摘要 AI 辅助
大语言模型(LLM)智能体在长程任务上仍不可靠,微小的局部故障会在长时间交互中累积,最终导致整体任务失败。尽管外部工具(harness)能大幅提升鲁棒性,但工具设计仍是手动且成本高昂的过程,需要在庞大的提示词、工具配置和控制逻辑空间中搜索。我们提出AutoSaddler,这是一个自动工具优化框架,将工具改进建模为离线学习问题,并使用小批量的故障信号迭代更新工具。AutoSaddler结合了故障轨迹诊断、将工具视为代码的结构化补丁生成,以及基于验证的更新选择。在GAIA2、SWE-Bench Pro和Terminal-Bench 2.0上的实验表明,AutoSaddler相比对应的基础工具大幅提升了智能体性能,分别实现了9.0、9.6和10.0个百分点的增益。消融研究进一步表明,有效的工具优化受益于三个关键要素:深度调试而非浅层反思、针对性修改而非无约束编辑、泛化感知选择而非轨迹特定修复。这些结果共同表明,自动工具优化是实现更高效、更可靠智能体系统的有前景路径。
英文摘要
LLM agents remain unreliable on long-horizon tasks, where small local failures can compound over extended interactions and lead to overall task failure. Although external harnesses can substantially improve robustness, harness design remains a manual and expensive process that requires searching over a large space of prompts, tool configurations, and control logic. We propose AutoSaddler, an automatic harness optimization framework that formulates harness improvement as an offline learning problem and iteratively updates the harness using failure signals from mini-batches. AutoSaddler combines failure-trace diagnosis, structured patch generation that treats the harness as code, and validation-based update selection. Experiments on GAIA2, SWE-Bench Pro, and Terminal-Bench 2.0 show that AutoSaddler substantially improves agent performance over the corresponding base harnesses, achieving gains of 9.0, 9.6, and 10.0 percentage points, respectively. Ablation studies further suggest that effective harness optimization benefits from three ingredients: deep debugging rather than shallow reflection, targeted modifications rather than unconstrained editing, and generalization-aware selection rather than trajectory-specific repair. Together, these results suggest that automatic harness optimization is a promising path toward more performant and reliable agent systems.
发表机构
- KAIST(韩国科学技术院)
- Southern University of Science and Technology(南方科技大学)
- Microsoft(微软)
- POSTECH(浦项科技大学)
机构由 AI 辅助整理,请以论文原文为准。