发表机构
Institute for Interdisciplinary Information Sciences Tsinghua University(清华大学交叉信息研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对LLM智能体的技能选择问题,提出了具备可证明双准则保证的BPS算法,在BigCodeBench变体上实现了更高任务成功率且标记使用更少的优势。
AI 中文摘要
将可复用的技能文档加载到有限的上下文窗口中,是当前大语言模型(LLM)智能体获取任务特定能力的主要方式,这使得技能选择成为决定任务性能和标记成本的首要因素。然而,现有的智能体仅通过语义相关性对技能进行独立评分,并通过Top-k或贪心打包来组装技能集,对所选技能集没有质量保证或成本意识。因此,冗余或选择不当的技能会浪费宝贵的上下文标记,甚至可能降低性能。我们首次建立了所选技能集如何影响执行结果的模型,并将技能选择转化为一个优化问题:在硬标记预算约束下选择技能集,以最大化单调子模效益减去上下文惩罚。针对该问题,我们开发了多项式时间算法Best Prefix Selection(BPS,最优前缀选择),并证明了据我们所知首个技能选择的性能保证:双准则$(1-1/e,1)$近似,其效益系数在多项式时间内是最优的。在经过污染控制的BigCodeBench变体上,BPS的表现优于所有基线方法,达到0.73的任务成功率,而已发布的技能路由器、文本检索器以及执行器自身选择的任务成功率为0.20至0.52,且BPS使用的标记比最强的已发布路由器少28%。
英文摘要
Loading reusable skill documents into a bounded context window has become a primary way large language model (LLM) agents acquire task-specific capabilities, which makes skill selection a first-order determinant of task performance and token cost. Yet current agents score skills independently by semantic relevance and assemble the set by top-$k$ or greedy packing, with no quality guarantee or cost awareness on the selected set. Redundant or poorly chosen skills then waste scarce context tokens and can even degrade performance. In this paper, we present a theory-grounded and practical framework for budgeted skill selection. We give the first model of how skill sets shape execution outcomes, capturing complementary capability coverage and diminishing returns from redundancy through a monotone submodular benefit, while accounting for context degradation with a linear token penalty under a hard budget. Based on this model, we develop Best Prefix Selection (BPS), a polynomial-time algorithm, and prove, to our knowledge, the first performance guarantee for skill selection: a bicriteria $(1-1/e,1)$ approximation whose benefit coefficient is optimal in polynomial time. We construct a controlled testbed based on BigCodeBench to isolate the effect of skill selection on execution success. On it, BPS with a learned capability encoder reaches a success rate of 0.65, and the strongest baselines need at least 28% more tokens to reach 0.60.