发表机构
Microsoft(微软)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究比较四种智能体技能加载方法,发现Hybrid等方法可显著减少token使用,且无质量损失,条件加载在技能大部分无需每轮使用时最有益。
AI 中文摘要
智能体技能通常在每次请求时完整注入,这会增加token成本。我们比较四种保留内容的加载方法:Full(完整加载)、Skill Block(技能块)、Reference(引用式)和Hybrid(混合式)。在SearchQA、SpreadsheetBench、ALFWorld、ScienceWorld和SynthProc数据集上,我们测量单轮任务的原始输入token使用量,以及多轮任务的缓存正确有效输入token使用量。结果显示没有通用最优方法:Hybrid在SearchQA上减少输入27.4%,在SpreadsheetBench上减少39.8%;在多轮长技能上,Skill Block和Hybrid实现大幅减少,在ScienceWorld上分别达62.5%和52.8%,在SynthProc上分别达73.0%和66.6%;ALFWorld的收益较小,因其流程简短且需重复使用。配对结果测试未检测到质量差异,但未确立等价性。总体而言,当技能的大部分在每轮都不需要时,条件加载最有益。
英文摘要
Agent skills are often injected in full on every request, increasing token cost. We compare four content-preserving loading methods: Full, Skill Block, Reference, and Hybrid. Across SearchQA, SpreadsheetBench, ALFWorld, ScienceWorld, and SynthProc, we measure token usage using raw input for single-turn tasks and cache-correct effective input for multi-turn tasks. Results show no universal winner. Hybrid reduces input by 27.4% on SearchQA and 39.8% on SpreadsheetBench. On large multi-turn skills, Skill Block and Hybrid achieve substantial reductions, reaching 62.5% and 52.8% on ScienceWorld and 73.0% and 66.6% on SynthProc. ALFWorld shows smaller gains because procedures are short and repeatedly needed. Paired outcome tests detect no quality differences, though they do not establish equivalence. Overall, conditional loading is most beneficial when large portions of a skill are not needed on every turn.