发表机构
Sichuan University(四川大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对良性技能组合可能引发恶意行为的问题,提出CRIME框架,通过分解目标、迭代精炼和沙箱验证,系统性地实例化组合攻击,并构建4000个技能的基准进行评估。
AI 中文摘要
智能体技能封装了特定任务的知识和程序,这些技能可以组合以支持复杂的智能体任务,而公共市场提供了日益增长的可复用技能池。然而,现有的安全审查大多孤立地评估技能,导致组合引发的风险未被充分探索。此类风险之所以产生,是因为组合良性技能扩展了智能体的能力空间,使得任何单一技能都无法实现的行为得以发生。有趣的是,我们发现直接组合良性技能即使每个技能都通过安全审查,也能诱导恶意行为。我们进一步发现,某些目标恶意行为即使所选技能共同提供了所需能力,也难以通过直接组合实现。为了系统性地实例化这些攻击,我们提出了多技能执行组合风险诱导(CRIME)方法。CRIME首先使用恶意情节编排(MPC)模块将目标恶意行为分解为互补的需求,并从公共技能库中识别合适的良性技能组合。对于无法直接实现目标行为的组合,失控反应引导(RRS)模块利用执行反馈迭代地精炼所选技能以趋近目标,同时要求每个技能在独立审查下保持良性。得到的组合随后传递给技能反应室(SRC)模块,在该模块中,技能对在沙箱中执行,并检查由此产生的环境后果以确定目标行为是否发生。未成功的案例返回RRS进行进一步精炼。此外,我们构建了一个包含跨八个网络安全行为的4,000个公共技能的基准,用于系统评估组合引发的漏洞。
英文摘要
Agent skills package task-specific knowledge and procedures that can be composed to support complex agent tasks, while public marketplaces provide a growing pool of reusable skills. Existing security vetting, however, largely evaluates skills in isolation, leaving composition-induced risks underexplored. Such risks arise because composing benign skills expands the agent's capability space, enabling behaviors unavailable to any skill alone. Interestingly, we find that directly composing benign skills can already induce malicious behaviors, even when every individual skill passes security vetting. We further find that some target malicious behaviors remain difficult to realize through direct composition, even when the selected skills collectively provide the required capabilities. To systematically instantiate these attacks, we present Compositional Risk Induction via Multi skill Execution (CRIME). CRIME first uses the Malicious Plot Casting (MPC) module to decompose a target malicious behavior into complementary requirements and identify suitable benign skill compositions from public skill repositories. For compositions that cannot directly realize the target behavior, the Runaway Reaction Steering (RRS) module uses execution feedback to iteratively refine the selected skills toward the target while requiring each skill to remain benign under standalone vetting. The resulting composition is then passed to the Skill Reaction Chamber (SRC) module, where the skill pair is executed in a sandbox and the resulting environmental consequences are examined to determine whether the target behavior has occurred. Unsuccessful cases are returned to RRS for further refinement. Furthermore, we construct a benchmark of 4,000 public skills across eight cybersecurity behaviors for systematic evaluation of composition-induced vulnerabilities.