arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SkillAtlas:智能体技能的攻击轨迹库

SkillAtlas: An Attack Trace Library for Agent Skills

Yuxin Tian, Zenghao Duan, Liang Pang, Zhiyi Yin, Xueqi Cheng

arXiv 2609.13353首次发表:更新:

发表机构

Institute of Computing Technology, CAS; University of Chinese Academy of Sciences(中国科学院计算技术研究所; 中国科学院大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SkillAtlas构建了一个包含3,014个案例、6,589条轨迹的攻击轨迹库,将私有安全报告转为可搜索公共案例,并利用轨迹标签将执行前防护准确率提升至0.770。

AI 中文摘要

智能体技能是语言模型智能体的可复用单元,但其风险通过模型决策、用户上下文、工具调用和执行反馈显现,而非通过稳定签名或单次沙箱运行。现有静态、动态和基准式评估很少保留可供检查、搜索和复用的公共证据。我们提出SkillAtlas,一个托管的攻击轨迹库,将私有的智能体技能安全报告包转换为经过审查、脱敏且可搜索的公共案例。该库包含3,014个案例、6,589条轨迹、151,131个步骤、233个受影响技能和8个风险类别;42.5%的成功案例在非成功的初始轮次后才首次成功,且基于轨迹的标签将执行前防护准确率提升至0.770。

英文摘要

Agent skills are reusable units for language-model agents, but their risks emerge through model decisions, user context, tool calls, and execution feedback rather than through stable signatures or a single sandbox run. Existing static, dynamic, and benchmark-style evaluations rarely preserve public evidence that can be inspected, searched, and reused. We present SkillAtlas, a hosted attack trace library that converts private agent-skill security report bundles into reviewed, redacted, and searchable public cases. The library contains 3,014 cases, 6,589 traces, 151,131 steps, 233 affected skills, and 8 risk categories; 42.5% of successful cases first become successful after a non-success initial round, and trajectory-grounded labels improve pre-execution guard accuracy to 0.770.

CommentsAccepted at REALM @ EMNLP 2026 (non-archival workshop paper). 9 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑