发表机构
Zhejiang University; Ant Group; Hangzhou Dianzi University; Tsinghua University(浙江大学; 蚂蚁集团; 杭州电子科技大学; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出首个多模态技能图像攻击安全基准MMSkillRisk,设计NCVA攻击,在九种配置中均诱导未授权操作,ASR达43.1%,表明任务成功不等于技能使用安全。
AI 中文摘要
智能体技能是可共享的过程性指令、工具和示例的封装包。多模态技能额外包含视觉参考,智能体在执行期间会检索并检查这些视觉参考。由于这些图像指导行动,攻击者可以将恶意指令伪装成看似合法的技能中的普通视觉指导。现有的技能安全研究主要考察文本携带的攻击或扫描器检测,对图像携带攻击在运行时的影响评估不足。我们引入了MMSkillRisk,据我们所知,这是第一个专门用于多模态技能中图像携带攻击的端到端安全评估的公开基准。为了实例化这一攻击面,我们设计了原生上下文视觉攻击(NCVA),该攻击将恶意指令伪装成教学图像的原生组件,如注释和界面标签。随附的此http URL提供指向相关视觉区域的辅助指导,而不明确陈述恶意操作。基于28个精心策划的干净技能,MMSkillRisk包含36个攻击包和108个可执行案例,涵盖五种攻击目标,并分别检查攻击成功和合法任务完成情况。在隔离沙箱中评估的九种模型-框架配置中,NCVA在每种配置中都诱导了未经授权的操作。其汇总攻击成功率(ASR)达到43.1%,比匹配的文本携带基线高出16.4个百分点,且在全部九种配置中ASR均更高。攻击成功与合法任务完成在36.5%的案例中同时发生,对于GPT-5.6-sol与Codex的组合,这一比例达到72.2%。这些结果表明,即使智能体完成合法任务,技能捆绑的图像也能诱导未经授权的操作,因此仅凭任务成功并不能确立技能使用的安全性。我们的代码和数据可在该https URL获取。
英文摘要
Agent skills are shareable packages of procedural instructions, tools, and examples. Multimodal skills additionally include visual references that agents retrieve and inspect during execution. Because these images guide actions, attackers can disguise malicious instructions as ordinary visual guidance within otherwise legitimate skills. Existing skill-security research primarily examines text-carried attacks or scanner detection, leaving the runtime effects of image-borne attacks insufficiently evaluated. We introduce MMSkillRisk, to our knowledge the first publicly available benchmark dedicated to end-to-end safety evaluation of image-borne attacks in multimodal skills. To instantiate this attack surface, we design Native-Context Visual Attack (NCVA), which disguises malicious instructions as native components of teaching images, such as annotations and interface labels. The accompanying SKILL.md provides auxiliary guidance toward relevant visual regions without explicitly stating the malicious operation. Built from 28 curated clean skills, MMSkillRisk contains 36 attack packages and 108 executable cases spanning five attack objectives, with separate checks for attack success and legitimate-task completion. Across nine model-harness configurations evaluated in isolated sandboxes, NCVA induces unauthorized operations in every configuration. Its pooled attack success rate (ASR) reaches 43.1%, exceeding the matched text-carrier baseline by 16.4 percentage points, with higher ASR in all nine configurations. Attack success and legitimate-task completion co-occur in 36.5% of cases, reaching 72.2% for GPT-5.6-sol with Codex. These results show that skill-bundled images can induce unauthorized actions even as agents complete legitimate tasks, so task success alone does not establish safe skill use. Our code and data are available at https://github.com/kaill-jlq/MMSkillRisk.