arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SkillEval:将智能体技能质量分解为可解释信号

SkillEval: Decomposing Agent Skill Quality into Interpretable Signals

Jiahui Han, Qinuo Li, Ziheng Peng, Haotian Wu, Haoze Liu, Danfeng Shan, Guanchu Wang, Huiqi Deng, Ninghao Liu

arXiv 2608.06891首次发表:更新:

AI 中文总结

本研究针对现有智能体技能评估仅反映任务适配性的不足,提出可解释框架SkillEval,其可区分技能质量、关联下游任务性能,还能指导技能修订以提升任务通过率。

AI 中文摘要

智能体技能提供可复用的程序性知识,助力智能体解决特定任务。随着其应用范围扩大,评估技能质量愈发重要。现有评估常通过测试技能是否提升特定下游任务的性能来衡量技能质量,但可复用技能适用于多个任务场景,下游评估仅反映技能与被评估任务的适配性,仅能提供技能质量的部分视角,无法确定技能应改进的方面。我们发现,技能文档(the this http URL document)的一般属性对技能质量起重要作用。为评估这些属性,我们提出SkillEval,一个用于文档级技能评估的可解释框架。SkillEval使用固定且可检查的评分方向评估每个属性,生成可解释的分数;还会测量并降低不相关文档特征(如长度和格式)的影响,使每个分数更具体地捕获其预期语义属性。具体而言,SkillEval从模型隐藏表示空间中的受控正负技能对中,为每个质量属性学习一个可解释方向,并通过将新技能的表示投影到这些固定方向来对其评分。我们使用SkillEval在受控质量测试中评估技能,结果显示SkillEval可可靠区分不同质量的技能;此外,SkillEval的评分与下游任务性能密切相关,可提前指示技能是否可能帮助智能体完成任务。我们进一步探索SkillEval用于诊断技能文档的弱点并指导针对性修订,修订后的技能改进了目标属性,在下游任务中实现了更高的通过率。

英文摘要

Agent skills provide reusable procedural knowledge that helps agents solve specialized tasks. As their use expands, evaluating skill quality becomes increasingly important. Existing evaluations often measure skill quality by testing whether a skill improves performance on specific downstream tasks. However, a reusable skill may apply to multiple task scenarios. Downstream evaluation mainly reflects the compatibility between a skill and the evaluated task, provides only a partial view of skill quality, and does not identify which aspect of the skill should be improved. We find that general properties of the \texttt{SKILL.md} document play an important role in skill quality. To evaluate these properties, we propose \textbf{SkillEval}, an interpretable framework for document-level skill evaluation. SkillEval evaluates each property using a fixed and inspectable scoring direction, producing interpretable scores. It further measures and reduces the influence of unrelated document features, such as length and formatting, so that each score captures its intended semantic property more specifically. Specifically, SkillEval learns an interpretable direction for each quality property from controlled positive--negative skill pairs in the hidden representation space of the model, and scores a new skill by projecting its representation onto these fixed directions. We use SkillEval to evaluate skills in controlled quality tests and show that SkillEval reliably distinguishes skills of different quality. In addition, SkillEval scores closely reflect downstream task performance, providing an early indication of whether a skill is likely to help an agent complete a task. We further explore SkillEval for diagnosing weaknesses in skill documents and guiding targeted revisions. The revised skills improve the targeted properties and achieve higher pass rates on downstream tasks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑