发表机构
Princeton University; UC San Diego; Stanford University; University of Southern California; Johns Hopkins University(普林斯顿大学; 加州大学圣迭戈分校; 斯坦福大学; 南加利福尼亚大学; 约翰斯·霍普金斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究探究 LLM 智能体技能的有效与失效原因,通过对比实验与轨迹分析揭示技能的作用机制,发现其主要靠程序锚定稳定动作,还指出检索瓶颈等问题,为智能体评估与优化提供指导。
AI 中文摘要
技能作为一种实用且有效的方法,通过结构化的知识包在推理时增强大型语言模型(LLM)智能体的能力。然而,现有评估大多仅衡量技能是否提升了聚合任务成功率,却忽视了一个更根本的问题:技能何时有帮助、为何有效、又在何处失效?本文通过在各类基准、智能体 harness 及 LLM 上开展的受控实验,分离出技能的表征、结果标注、检索难度及跨框架鲁棒性的影响。为进一步解答该问题,本文设计了一项对比研究,结合受控定量实验与配对轨迹分析,对受控实验中的 8135 条试验记录进行归一化处理,从 240 条开放编码记录中保留 238 个有效唯一标签,并将这些观察结果整合为包含三个高级类别和十二种技能使用模式的分类体系。技能在噪声轨迹成为稳定执行的程序锚点时有效;在匹配对比中,技能较工作流记忆(Workflow Memory)提升了 6.06 个百分点;程序锚定占技能案例的 65.7%,而显式知识注入仅占 4.5%,表明技能是通过稳定动作而非注入缺失事实发挥作用。检索是一个独立瓶颈:当检索池从 5 增长到 100 时,实际使用准确率从 29.6% 降至 3.3%;易混淆的干扰项会损害离线识别,但下游成功率仍保持稳定,准确调用真实值既非充分也非必要条件。技能在假设脆弱、上下文不兼容或适配不足时失效。这些发现将评估范围从聚合成功率拓展,并为可靠的自进化智能体提供指导。
英文摘要
Skills have emerged as a practical and effective approach for enhancing LLM agents at inference time through structured packages of knowledge. However, existing evaluations largely measure whether skills improve aggregated task success, leaving a more fundamental question underexplored: \emph{\textbf{When do skills help, why do they work, and where do they fail?}} Through controlled experiments across various benchmarks, agent harnesses and LLMs, we isolate the effects of representation, outcome annotation, retrieval difficulty, and cross-framework robustness of skills. To further answer this question, we design a contrastive study that combines controlled quantitative experiments with paired trajectory analysis. We normalize 8,135 trial records from controlled experiments and retain 238 valid unique labels from 240 open-coded records. We consolidate these observations into a taxonomy of three high-level categories and twelve skill-use modes: skills work when noisy trajectories become procedural anchors that stabilize execution. Skills improve over Workflow Memory by 6.06 points in matched comparisons. Procedural anchoring accounts for 65.7\% of skill cases, versus 4.5\% for explicit knowledge injection, showing that skills stabilize action rather than inject missing facts. Retrieval is a separate bottleneck: as pools grow from 5 to 100, actual-use precision falls from 29.6\% to 3.3\%. Confusable distractors impair offline identification, yet downstream success remains stable; exact ground-truth invocation is neither sufficient nor necessary. Skills fail under brittle assumptions, incompatible contexts, or insufficient adaptation. These findings move evaluation beyond aggregate success rates and guide reliable self-evolving agents.