发表机构
University of California, Santa Barbara; MIT-IBM Watson AI Lab; Adobe Research(加州大学圣塔芭芭拉分校; MIT-IBM沃森人工智能实验室; Adobe研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出SHarP,一种基于模块显著性的智能体框架剪枝方法,通过消融评估各模块对性能和token成本的影响,发现多数框架高度冗余,剪枝大量模块后仍能保持可比性能与效率。
AI 中文摘要
智能体框架(harness)是协调模型调用、工具使用和任务执行以帮助大型语言模型完成复杂任务的系统。为满足任务需求并处理失败情况,这些系统通常通过修改和修补其指令、工具和工作流来迭代优化,从而不断增加框架的复杂性。因此,尚不清楚某些生成的框架模块是否冗余,是否引入了大量token开销而几乎没有性能提升。受神经网络剪枝的启发,本文研究框架剪枝作为在任务性能和token成本之间取得更好平衡的手段。我们提出SHarP(基于显著性的框架剪枝),一种简单而有效的剪枝策略,基于每个框架模块相对于性能和效率的显著性。具体而言,我们首先将工具、指令和支持机制识别为可单独禁用的组件。然后,通过消融每个模块并评估其相对于完整单模块消融集的任务性能和token成本来估计其显著性。贡献最小或计算开销最大的模块随后被剪枝。我们在保留验证集上对各种框架的评估揭示了一个令人惊讶的发现:我们研究的大多数框架高度冗余,即使在剪枝掉相当大比例的模块后,仍能保持可比的性能和效率。我们的剪枝方法和实证发现为智能体框架的设计和优化提供了新视角。
英文摘要
Agent harnesses are systems that coordinate model calls, tool use, and task execution to help large language models complete complex tasks. To meet task requirements and address failures, these systems are often iteratively refined by amending and patching their instructions, tools, and workflows, continuously increasing harness complexity. It is therefore unclear whether some resulting harness modules are redundant, introducing substantial token overhead with little, if any, performance gain. Inspired by neural network pruning, in this paper, we study harness pruning as a means of striking a better balance between task performance and token cost. We propose SHarP (Saliency-based Harness Pruning), a simple yet effective pruning strategy based on the saliency of each harness module with respect to performance and efficiency. Specifically, we first identify tools, instructions, and supporting mechanisms as components that can be individually disabled. We then estimate the saliency of each module by ablating it and assessing its task performance and token cost relative to the full set of single-module ablations. Modules with the smallest contribution to performance or largest computational overhead are subsequently pruned. Our evaluation across various harnesses on held-out validation sets reveals a surprising finding: most harnesses that we studied are highly redundant and can maintain comparable performance and efficiency even after a substantial portion of their modules are pruned. Our pruning approach and empirical findings provide new perspectives on agent harness design and optimization.