arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.19901cs.CRcs.AI

MaliciousSkillBench:用于恶意智能体技能检测的综合基准

MaliciousSkillBench: A Comprehensive Benchmark for Malicious Agent Skill Detection

Yue Wang, Yi Liu, Gelei Deng, Ying Zhang, Yuekang Li, Zhenyu Chen, Leo Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出MaliciousSkillBench这一综合基准,整合多源恶意技能数据构建含9740个技能的数据集,评估三类检测器后发现现有方法不足,需更广泛跨源覆盖及联合评估攻击检测与良性误报的方案。

中文摘要 AI 辅助

智能体技能为大语言模型(LLM)智能体提供可复用的指令包,其中可能包含脚本、资源和服务配置,这为恶意行为创造了直接的传播渠道。然而,现有的恶意技能数据集在来源、制品格式、证据体系和良性覆盖范围上存在碎片化问题,重复及结构相关的内容进一步增加了直接聚合与评估的难度。本文提出MaliciousSkillBench,这是一个用于恶意智能体技能检测的综合基准。我们整合了13个公开来源,其中11个提供核心恶意制品,将8414条原始恶意记录精简为4588个操作结构家族中的7539个标准化唯一身份。经过保守的跨标签冲突排除后,主基准包含9740个技能:7505个恶意技能和2235个良性技能。为表征其覆盖范围,我们协调了11个攻击类别,对应4983个恶意身份及支持的来源原生映射,并发现不同来源的威胁构成存在显著差异。随后,我们评估了3个学习型文本检测器和3个现成的技能扫描器。学习型检测器在随机宏F1值上达到0.882-0.932,但在来源不相交评估下仅为0.653-0.665;最强的词TF-IDF SVM在随机/结构不相交/来源不相交评估中分别得分为0.932/0.916/0.665,同时保持95.6%的恶意召回率,但在保留来源上产生62.4%的良性假阳性率。现成扫描器占据不同但同样不令人满意的操作区间,其降低假阳性的代价是恶意召回率大幅下降。这些结果共同表明,可靠的恶意技能检测需要更广泛的跨源基准覆盖,以及同时衡量攻击检测和良性误报的评估方式。

英文摘要

Agent Skills extend LLM agents with reusable instruction packages that may also include scripts, resources, and service configuration. This creates a direct distribution channel for malicious behavior, yet existing malicious-Skill datasets are fragmented across sources, artifact formats, evidence regimes, and benign coverage; duplicated and structurally related content further complicates direct aggregation and evaluation. We present MaliciousSkillBench, a comprehensive benchmark for malicious Agent Skill detection. We consolidate 13 public sources, 11 of which contribute Core malicious artifacts, and reduce 8,414 raw malicious records to 7,539 normalized-unique identities in 4,588 operational structural families. After conservative cross-label conflict exclusion, the primary benchmark contains 9,740 Skills: 7,505 malicious and 2,235 benign. To characterize its coverage, we harmonize 11 attack categories for 4,983 malicious identities with supported source-native mappings and find substantial differences in threat composition across sources. We then evaluate three learned text detectors and three off-the-shelf Skill scanners. Learned detectors achieve 0.882-0.932 Random Macro-F1 but only 0.653-0.665 under Source-Disjoint evaluation; the strongest word TF-IDF SVM scores 0.932/0.916/0.665 on Random/structural-disjoint/Source-Disjoint while retaining 95.6% malicious recall but producing 62.4% benign FPR on held-out sources. Off-the-shelf scanners occupy different but also unsatisfactory operating regimes, reducing false positives only at the cost of sharply lower malicious recall. Together, these results show that reliable malicious-Skill detection requires both broader cross-source benchmark coverage and evaluation that jointly measures attack detection and benign over-flagging.

发表机构

  • Nanjing University(南京大学)
  • Griffith University(格里菲斯大学)
  • Nanyang Technological University(南洋理工大学)
  • Wake Forest University(维克森林大学)
  • University of New South Wales(新南威尔士大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑