发表机构
Central China Normal University; Beijing Kuafu Technology Co., Ltd.; Monash University(华中师范大学; 北京夸父科技有限公司; 蒙纳士大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SkillApt提出检索后激活框架,基于反事实证据决定是否加载技能,在保持准确率的同时大幅降低激活率和token使用量。
AI 中文摘要
大型语言模型智能体越来越多地检索可复用的技能并将其注入到活动上下文中。然而,检索到的技能在当前执行状态下可能相关但不必要、代价高昂,甚至有害。我们提出了SkillApt,一个检索后激活框架,用于决定检索到的技能是否应被实际加载。SkillApt从匹配的WITH/WITHOUT运行中构建执行证据,并利用相似历史状态的结果为每个候选技能做出LOAD/ABSTAIN(加载/弃权(不执行))决策。在冻结的确认性SRA-Bench评估中,SkillApt-E实现了与BM25 Top-1相同的观测准确率(0.838对0.838),同时将技能激活率从100%降低至31.5%,平均token使用量减少了74.3%。进一步的诊断表明,技能效用及其激活边界的可学习性因基础模型而异。这些结果表明,技能检索和技能激活应被视为独立的决策:检索识别哪些技能可能相关,而SkillApt则确定在当前状态下使用该技能是否值得。
英文摘要
Long-running agents accumulate reusable Skills, but a Skill that is semantically relevant to a task is not necessarily worth loading in the current state. We study the applicability question that arises once a candidate Skill is known: should it be loaded in the current state? We propose SkillApt, which uses matched WITH/WITHOUT Skill executions on the same task state as persistent evidence, estimates the conditional marginal utility of the Skill, and chooses LOAD or ABSTAIN accordingly. The base model, agent architecture, and Skill contents stay fixed; only the external deployment policy changes. We call this constrained setting strategy-level recursive self-improvement (Strategy-Level RSI). On 20 Skills and 160 held-out states, as paired evidence accumulates, SkillApt's task success rises from 81.9% under a cold start to 91.3%, matching a strong zero-shot LLM controller; yet SkillApt activates Skills on only 26.3% of states, versus 98.8% for the zero-shot controller. A hard-candidate study shows that non-optimal Skills mostly leave correctness unchanged while raising execution cost, and occasionally cause correctness harm. An ablation shows that a history recording only WITH success makes the policy load almost everywhere, whereas paired evidence substantially improves selectivity. These results indicate that relevance is not applicability: the main effect of execution evidence is not to make the model stronger but to change how existing Skills are deployed, moving the system from near-always loading to selective reuse.
Comments21 pages, 7 figures, 13 tables. Code: https://github.com/guoshaung/SkillApt ; Project page: https://gym.codeflying.net/