AI 中文总结
SkillConsist针对智能体技能不一致检测的挑战,通过双向图对齐等技术构建基准并取得优于基线的检测性能,提升了技能一致性检测效果。
AI 中文摘要
智能体技能为大语言模型(LLM)智能体提供可复用的能力,其不一致性可能会暴露未公开的危险行为或导致技能选择错误。近期智能体技能研究日益关注技能一致性检测,现有方法会针对预定义类别或声明范围评估行为或安全属性图,而PL-HCL方法采用基于LLM的模型学习元数据、指令与资源间的一致性。但声明与实现行为可能混杂在文本和代码中,且简洁声明可对应多个关联实现步骤,为此本文提出SkillConsist解决这两个挑战:LLM将声明和实现内容分别拆分为声明侧与实现侧的行为记录,静态分析补充实现侧记录,这些记录分别构成声明行为图与实现行为图;从任意一侧的行为记录出发,双向图对齐会在另一张图中搜索候选子图,并沿行为关系扩展直至完整表达源侧行为,图差异识别对齐子图间的冲突并输出检测结果。本文从ClawHub的500个下载量最高的公开技能及133个Skill-Inject包构建633-Skill基准,该基准含319个不一致技能、314个一致技能及442个定位的不一致标注;在该基准上,SkillConsist在包级检测中达到86.85%的精确率、89.03%的召回率、87.93%的F1值,较最优基线的F1值提升20.43个百分点,在定位任务中达到67.60%的精确率、58.14%的召回率、62.52%的F1值。
英文摘要
Agent Skills provide reusable capabilities to LLM agents. Agent Skill inconsistencies can expose undisclosed dangerous behavior or cause wrong Skill selection. Recent Agent Skill research has increasingly examined Agent Skill consistency detection. Existing methods evaluate behaviors or security-property graphs against predefined categories or declared scopes. More recently, PL-HCL uses an LLM-based model to learn consistency across metadata, instructions, and resources. However, declaration and implementation behavior can be mixed across text and code, and a concise declaration can correspond to multiple connected implementation steps. We present SkillConsist to address both challenges. An LLM separates declaration and implementation content into behavior records on the implementation and declaration sides, while static analysis supplements implementation records. These records form declaration and implementation behavior graphs, respectively. Starting from a behavior record on either side, bidirectional graph alignment searches the other graph for a candidate subgraph and expands it along behavior relations until it completely expresses the source-side behavior. Graph differencing identifies conflicts between aligned subgraphs and outputs the detection results. We construct a 633-Skill benchmark from ClawHub's 500 most-downloaded public Skills and 133 Skill-Inject packages. The benchmark contains 319 inconsistent and 314 consistent Skills and 442 localized inconsistency annotations. On this benchmark, SkillConsist achieves 86.85% precision, 89.03% recall, and 87.93% F1 for package-level detection, improving F1 over the best baseline by 20.43 percentage points. For localization, it achieves 67.60% precision, 58.14% recall, and 62.52% F1.
Comments11 pages, 3 figures