AI 中文总结
研究基于语言模型的智能体使用第三方危险技能时的安全问题,构建OpenSkillRisk基准测试,涵盖多种不安全场景并能细粒度分析。实验发现当前系统处理危险技能能力不足,揭示三种失败模式,强调需改进语言模型风险推理和智能体框架执行控制。
AI 中文摘要
基于语言模型的智能体在开放世界场景中利用第三方技能扩展其能力。然而,第三方技能可能会引入额外的安全漏洞,看似无害的技能可能包含潜在安全风险,仅在实际执行时出现。本文系统研究了当前智能体系统识别和避免此类风险的能力。为支持定量和定性评估,构建了OpenSkillRisk基准测试,包含从公共技能市场收集的263个危险技能,按威胁类型分为七类,并为每个技能配备标准化用户任务和相应沙盒进行评估。与先前基准测试不同,它不仅涵盖更现实多样的不安全场景,还提供细粒度分析。对三个主流命令行智能体框架和13个先进语言模型进行综合实验,结果表明无测试系统能可靠处理危险技能,即使最安全配置仍有17%的情况执行不安全操作。上下文相关和系统级风险尤其难以避免,行为分析揭示了三种常见失败模式。这些发现凸显了改进语言模型风险推理和智能体框架执行控制的必要性。
英文摘要
LLM-based agents leverage third-party skills to extend their capabilities in open-world scenarios. However, third-party skills can introduce extra security vulnerabilities, as seemingly harmless skills can contain latent safety risks that only emerge during actual execution. In this work, we conduct a systematic investigation into how well current agent systems recognize and avoid such risks. To support quantitative and qualitative evaluation, we construct OpenSkillRisk, a dedicated safety benchmark containing 263 risky skills collected from public skill marketplaces. We classify these skills into seven categories based on their threat types and pair each skill with a standardized user task and a corresponding sandbox for controlled evaluation. Distinct from prior benchmarks, OpenSkillRisk not only covers more realistic and diverse unsafe scenarios, but also provides a fine-grained analysis to diagnose the behavioral patterns of agents in such scenarios. We conduct comprehensive experiments covering three mainstream CLI agent frameworks and thirteen state-of-the-art LLMs. Experimental results show that no tested system handles risky skills reliably: even the safest configurations still execute unsafe actions in about 17% of cases. Context-dependent and system-level risks are especially difficult for current agent systems to avoid. Our behavioral analysis reveals three recurring failure patterns: agents may fail to recognize the risk, recognize it but fail to intervene before acting, or follow skill instructions beyond the user's intended scope. These findings highlight the need to improve both risk reasoning in LLMs and execution control in agent frameworks.