技能使用:智能体框架中的大语言模型(LLM)真的能使用技能吗?
Skill-Use: Can LLMs Actually Use Skills in Agentic Harnesses?
浏览论文内容
中文总结 AI 辅助
本研究推出Skill-Use基准,评估智能体在渐进式披露下的技能使用,发现8个LLM在两个框架下的SU分数仅0.613,技能使用是依赖框架的能力。
中文摘要 AI 辅助
大语言模型(LLM)智能体越来越依赖技能,技能是指定何时行动、遵循何种流程以及允许使用哪些工具的结构化文档。现有评估大多判断技能的质量或其对任务成功的贡献,却未考察智能体能否自行识别相关技能并应用它。我们推出Skill-Use,这是一个在渐进式披露下评估技能使用的基准,其中智能体仅能看到技能的名称和简短描述,必须检索完整流程后再遵循。Skill-Use区分了技能使用的三个方面:触发(衡量智能体是否调用相关技能)、合规性(衡量其遵循规定流程的忠实程度)、边界(衡量其是否避免禁止操作)。Skill-Use(SU)分数结合了这三个方面,且仅在触发技能后才对执行情况计分。Skill-Use将79个真实技能与9个领域的177个可执行任务配对,每个任务都基于真实文件,在隔离的Docker沙箱中运行,并通过基于轨迹的评分标准计分。在两个智能体框架下评估8个LLM后,我们发现可靠的技能使用仍遥不可及,因为最强配置的SU分数仅为0.613。触发和流程合规性是独立的瓶颈,且分数和模型排名会随框架变化,因此技能使用是一种依赖于框架而非模型固定属性的能力。
英文摘要
Large language model (LLM) agents increasingly rely on skills, structured documents that specify when to act, which procedure to follow, and which tools are allowed. Existing evaluations mostly judge the quality of a skill or its contribution to task success, leaving unexamined whether an agent can recognize a relevant skill and apply it on its own. We introduce Skill-Use, a benchmark that evaluates skill use under progressive disclosure, where an agent sees only a skill's name and short description and must retrieve the full procedure before following it. Skill-Use separates three facets of skill use. Trigger measures whether the agent invokes the relevant skill, Compliance measures how faithfully it follows the prescribed procedure, and Boundary measures whether it avoids forbidden operations. A Skill-Use (SU) score combines the three and credits execution only after the skill is triggered. Skill-Use pairs 79 real skills with 177 executable tasks across nine domains, each grounded in real files, run in an isolated Docker sandbox, and scored by a trajectory-based rubric. Evaluating eight LLMs under two agent harnesses, we find that reliable skill use remains out of reach, as the strongest configuration reaches an SU of only 0.613. Triggering and procedural compliance fail as independent bottlenecks, and both scores and model rankings shift with the harness, so skill use behaves as a capability conditioned on the harness rather than a fixed property of the model.
发表机构
- East China Normal University(华东师范大学)
- Hong Kong University of Science and Technology(香港科技大学)
- Fudan University(复旦大学)
- Tencent Hunyuan(腾讯混元)
机构由 AI 辅助整理,请以论文原文为准。