发表机构
Peking University; Zhongguancun Academy; Hefei University of Technology; Tsinghua University(北京大学; 中关村学院; 合肥工业大学; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对技能可能无益甚至有害的问题,提出SkillDelta框架,通过配对执行预测任务条件技能增益,在15个设置中平均提升成功率4.3%,并揭示增益主要来自跨任务组分配。
AI 中文摘要
智能体技能预期能提升任务性能。然而,我们发现技能往往无法带来益处,甚至可能损害性能,同时产生额外的令牌成本。我们能否在智能体行动之前预测技能是否有帮助?我们提出了SkillDelta,一个从同一智能体在有技能和无技能情况下的配对执行中预测任务条件技能增益的框架。一个局部预测器将这些历史增益迁移到新任务,而无需重新训练智能体。在显式迁移假设下,我们的分析将支持覆盖、表示不匹配和执行噪声与预测误差及决策遗憾联系起来。在五个基准和三个目标智能体上,配对历史在15个设置中的12个中改善了相对于仅技能辅助结果的观测增益排序。在匹配的预期技能使用率下,SkillDelta在所有15个设置中相对于随机激活提高了成功率,平均绝对增益为4.3%。这一优势大部分来自跨任务组的分配。额外的组内选择价值的证据在ToolQA上最强,在其他地方较弱。代码可在以下网址获取:此https URL。
英文摘要
Agent skills are expected to improve task performance. Yet we find that they often provide no benefit, and can even hurt performance while incurring additional token costs. Can we predict whether a skill will help before the agent acts? We introduce SkillDelta, a framework for predicting task-conditional skill gains from paired executions of the same agent with and without the skill. A local predictor transfers these historical gains to new tasks without retraining the agent. Under explicit transfer assumptions, our analysis links support coverage, representation mismatch, and execution noise to prediction error and decision regret. Across five benchmarks and three target agents, paired history improves observed-gain ranking over skill-assisted outcomes alone in 12 of 15 settings. At matched expected skill-use rates, SkillDelta improves success over random activation in all 15 settings, with an average absolute gain of 4.3%. Most of this advantage comes from allocation across task groups. Evidence for additional within-group selection value is strongest on ToolQA and weaker elsewhere. Code is available at https://github.com/TankTechnology/skilldelta.
Comments22 pages, 7 figures. Code: https://github.com/TankTechnology/skilldelta