发表机构
University of Illinois Urbana-Champaign; University of Michigan, Ann Arbor(伊利诺伊大学厄巴纳-香槟分校; 密歇根大学安娜堡分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对固定智能体框架在不同任务上的次优性,提出基于可复用原语的STITCH框架,在测试时组合任务特定框架,提升成功率最多12个百分点,且开销极低。
AI 中文摘要
智能体框架(Agent harnesses)控制着大型语言模型(LLMs)如何收集上下文、调用工具、验证结果、保持状态以及终止运行,在很大程度上影响智能体的性能。然而,每种框架机制在不同任务上的价值可能不同:一种机制可能提升某个任务的性能,却对另一个任务造成额外开销或上下文干扰,从而导致全局框架的次优性。我们将这种次优性归因于固定机制选择所引发的不匹配,这促使我们构建任务特定的框架。尽管如此,为每个任务生成框架代码会带来生成和调试成本,且随着生成的机制增多,执行风险可能累积。为了应对这些挑战,我们引入了框架原语(Harness Primitives),即具有明确应用范围和组合契约的可复用框架机制,这些机制从失败的任务轨迹中挖掘而来。基于框架原语,我们提出了STITCH框架,该框架根据任务信息选择合适原语,并在测试时将它们编译为任务特定的框架。这种分离使得无需在测试时生成或修复机制代码即可实现任务特定框架。大量实验表明,STITCH不仅提升了框架的适应性和鲁棒性,而且随着原语库规模的扩大而扩展,相比固定框架基线,任务成功率提升了最多12个百分点,超越了如Codex CLI等人力设计的框架,同时保持了极低的测试时框架组合开销,仅为2.7%,比从头生成任务特定框架高效638倍。最终,我们的工作表明,构建任务自适应框架有助于完成多样化任务,而构建可复用原语是实现这一目标的有前景的途径。
英文摘要
Agent harnesses govern how large language models (LLMs) gather context, invoke tools, verify results, preserve state, and terminate, largely affecting agent performance. However, the value of each harness mechanism can differ across heterogeneous tasks: a mechanism that improves one task may impose overhead or context distraction on another, leading to the suboptimality of a global harness. We characterize this suboptimality as a mismatch induced by fixed mechanism choices, motivating task-specific harness construction. Nonetheless, generating harness code for each task introduces generation and debugging costs, with execution risks that can compound as more mechanisms are generated. To address those challenges, we introduce Harness Primitives, reusable harness mechanisms with clear application scope and composition contract mined from failed task trajectories. Based on Harness Primitives, we propose STITCH, a framework that Selects suitable primitives given Task Information and compiles them into Task-speCific Harnesses at test time. This separation enables task-specific harnesses without generating or repairing mechanism code at test time. Extensive experiments demonstrate that STITCH not only improves harness adaptability and robustness, but also scales with the primitive library size, boosting task success rates by up to 12 points over fixed harness baselines, surpassing human-designed harnesses like Codex CLI while maintaining a minimal test-time harness composition overhead of only 2.7%, 638 times more efficient than generating task-specific harnesses from scratch. Ultimately, our work demonstrates that building task-adaptive harnesses can be beneficial for completing diverse tasks and that building reusable primitives can be a promising path towards this goal.