arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向编码智能体中恶意技能文件的风险评估

Towards a Risk Assessment of Malicious Skill Files in Coding Agents

Rui Yang, Michael Fu, Kla Tantithamthavorn, Chetan Arora, Joey Chua

arXiv 2608.05223首次发表:更新:

AI 中文总结

该研究针对编码智能体技能接口的恶意文件风险,提出对抗性技能合成方法与评估流水线,发现企业级智能体易被恶意技能利用,仅少数运行有明确安全识别,呼吁企业提前评估风险。

AI 中文摘要

自主编码智能体正越来越多地嵌入企业软件工作流中,并拥有对连接系统的委托权限。该架构的核心是智能体技能接口:智能体动态加载以专门化其行为的指令和脚本文件夹。该接口也扩大了攻击面,让恶意 shell 命令隐藏在自然语言技能文件中。我们做出三项贡献:第一,一种对抗性技能合成方法,使用四个系列的六个大语言模型(LLM)将 471 个真实世界 shell 命令转换为外观良性的技能,发布为包含 2826 个技能的基准,这些技能映射到 11 种 MITRE ATT&CK 战术;第二,一种可复现的评估流水线,结合运行分层、证据锚定、弃权(不执行)否决机制和确定性声明意图覆盖,以及一个由三个评判者组成的 LLM-as-a-judge 小组,针对盲人类黄金标准进行验证(Cohen's kappa = 0.85);第三,对两个企业级智能体的大规模特征分析,涉及 5629 次完成的运行。Gemini CLI 在 95.5-96.1% 的运行中被利用,Qwen Code 在 71.6-74.0% 的运行中被利用(从原始多数投票到经声明意图修正的估计,均在人类黄金标准范围内),几乎不受生成模型影响。仅 1.99% 的运行出现明确的安全识别。企业在采用编码智能体前必须评估并缓解技能接口风险。我们的代码和数据集可在此 https URL 获取。

英文摘要

Autonomous coding agents are increasingly embedded in enterprise software workflows with delegated authority over connected systems. Central to this architecture is the agent skills interface: folders of instructions and scripts that agents load dynamically to specialize their behavior. This interface also widens the attack surface, letting malicious shell commands hide within natural-language skill files. We make three contributions. First, an adversarial skill-synthesis method using six LLMs across four families to transform 471 real-world shell commands into benign-appearing skills, released as a benchmark of 2,826 skills mapped to 11 MITRE ATT&CK tactics. Second, a reproducible evaluation pipeline coupling run stratification, evidence anchoring, a refusal veto, and a deterministic declared-intent override with a three-judge LLM-as-a-judge panel, validated against a blind human gold standard (Cohen's kappa = 0.85). Third, a large-scale characterization of two enterprise-grade agents across 5,629 completed runs. Gemini CLI is exploited in 95.5-96.1% of runs and Qwen Code in 71.6-74.0% (raw majority vote to declared-intent-corrected estimate, both within the human gold standard), nearly invariant to the generating model. Explicit safety recognition occurs in only 1.99% of runs. Enterprises must assess and mitigate skill-interface risk before adopting coding agents. Our code and dataset are available at https://github.com/awsm-research/AgentJailbreak

Comments29 pages, 6 figures, 6 tables. Preprint; under review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑