arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.08468cs.CRcs.AI

SkillsMetric:映射恶意智能体技能静态分析的检测边界

SkillsMetric: Mapping the Detection Boundary of Static Analysis for Malicious Agent Skills

  • The Graduate Center, City University of New York(纽约城市大学研究生中心)
  • Hunter College, City University of New York(纽约城市大学亨特学院)
  • The Ohio State University(俄亥俄州立大学)

机构由 AI 辅助整理,请以论文原文为准。

Xinze Chen, Chi Zhang, Ping Ji, Yimin Liu

AI总结:

本研究提出五阶段静态分析框架SkillsMetric,构建含2266个技能的对抗性数据集,发现静态分析对部分攻击检测不足,需结合静态预筛选与语义审查的纵深防御架构。

AI中文摘要:

基于大语言模型的智能体所使用的技能——用于增强智能体的结构化指令包与脚本——正快速普及,但其安全特性仍未得到充分探索。本文提出\textsc{SkillsMetric},这是一个五阶段静态分析框架,可从模式密度、统计异常、数据流污点、导入异常、能力不匹配五个维度对技能包进行评分。我们构建了包含2266个技能的对抗性评估数据集,涵盖代码级、系统级、语义级威胁的16种攻击类型,并在完整的SkillMD-138K语料库上进行评估。该框架的AUC达到0.93,五折交叉验证的F1值为73.4%±0.5%,对数据外泄(93%)和隐写 payload(93%)的检测能力较强。关键的是,我们发现了根本性的盲点:使用常见Shell命令的“主机破坏”攻击可规避全部五个阶段(检测率为0%),而通过自然语言操纵的“提示注入”攻击的检测率仅为42%。这些发现表明,仅靠静态分析不足以保障技能安全,因此需要构建纵深防御架构,将快速静态预筛选与语义审查相结合。

英文摘要:

Agent Skills---structured packages of instructions and scripts that augment LLM-based agents---are rapidly proliferating, yet their security properties remain under-explored. We present \textsc{SkillsMetric}, a five-stage static analysis framework that scores skill packages along pattern density, statistical anomaly, dataflow taint, import anomaly, and capability mismatch dimensions. We construct an adversarial evaluation dataset of 2{,}266 skills spanning 16~attack types across code-level, system-level, and semantic-level threats, and evaluate on the full SkillMD-138K corpus. Our framework achieves an AUC of 0.93 and 5-fold cross-validated F1 of 73.4\%$\pm$0.5\%, with strong detection of data exfiltration (93\%) and steganographic payloads (93\%). Crucially, we identify fundamental blind spots: \emph{host destruction} attacks using common shell commands evade all five stages (0\% detection), and \emph{prompt injection} via natural-language manipulation achieves only 42\% detection. These findings establish that static analysis alone is insufficient for skill security, motivating defense-in-depth architectures that combine fast static pre-screening with semantic review.

补充信息

↑