AI 中文总结
SkillShield是一种系统提示防御机制,通过离线合成安全技能注入系统提示,可有效降低LLM编码智能体的恶意软件生成成功率,在多项实验中表现优于基线方法。
AI 中文摘要
编码智能体以开发者权限编辑文件并执行Shell命令,使得恶意请求可直接转化为有害操作或功能性恶意软件。现有防御方法存在互补性局限:仅使用API的部署者无法应用权重级对齐,而输入过滤器和执行边界监控器需要在智能体的执行路径中添加辅助分类或检查组件。因此,我们提出SkillShield,一种系统提示防御机制,该机制可从已知攻击或记录的智能体失败案例中离线合成安全技能,并在会话开始时将这些技能注入系统提示,使其在整个工具使用循环中保持激活状态。与参考监控器不同,SkillShield通过定义模型在执行过程中应遵循的安全策略来保护系统。由于系统提示空间有限,我们研究了三种固定预算的配置范围:全类别范围(一种技能覆盖所有威胁类别)、按捆绑包范围(一种技能针对相关子集)、按类别范围(一种技能专注于单个已知类别,用作上限参考),所有配置均无需运行时请求分类或路由。在RedCode数据集上针对六个大语言模型的实验显示,默认的全类别技能将恶意软件生成严重程度从3.37降至0.58,执行攻击成功率达到43.6%,与Llama Guard 3的42.7%相当(Llama Guard 3未使用其单独的8B分类器);按捆绑包和按类别配置进一步将该成功率分别降至36.2%和14.5%。在两种非自适应越狱系列攻击下,SkillShield在恶意软件生成任务上仍优于所有基线方法。在731个良性任务描述中,SkillShield的平均安全弃权(不执行)率为0.14%。这些结果表明,提示空间安全技能具有防止LLM编码智能体执行有害操作和生成恶意软件的潜力。
英文摘要
A coding agent edits files and executes shell commands with its developer's privileges, allowing malicious requests to translate directly into harmful actions or functional malware. Existing defenses have complementary limitations: weight-level alignment is unavailable to API-only deployers, whereas input filters and execution-boundary monitors require auxiliary classification or checking components along the agent's trajectory. We therefore introduce SkillShield, a system-prompt defense that synthesizes security skills offline from known attacks or recorded agent failures. These skills are injected into the system prompt at session start and remain active throughout the tool-use loop. Unlike a reference monitor, they protect the system by defining the security policies the model should follow during execution. Due to the limited system-prompt space, we examine three fixed-budget provisioning scopes: all-classes, with one skill covering all threat classes, per-bundle, with one skill targeting a related subset, and per-class, with one skill dedicated to a single known class and used as the upper-bound reference. None requires runtime request classification or routing. Across six large language models on RedCode, the default all-classes skill reduces malware-generation severity from 3.37 to 0.58 and achieves a 43.6% execution attack success rate, comparable to Llama Guard 3's 42.7% without its separate 8B classifier. The per-bundle and class-fixed per-class settings further reduce this rate to 36.2% and 14.5%, respectively. Under two non-adaptive jailbreak families, SkillShield continues to outperform all baselines on malware generation. Across 731 benign task descriptions, SkillShield yields a mean safety-refusal rate of 0.14%. These results demonstrate the potential of prompt-space security skills to prevent harmful actions and malware generation for LLM coding agents.