AI 中文总结
SkillSentry是基于自适应蜜罐世界的动态安全测试框架,可测试大语言模型智能体外部技能的条件性有害行为,在基准测试及规避场景下性能优于基线,代码公开。
AI 中文摘要
外部技能扩展了大语言模型智能体的能力,但也引入了执行时攻击面:一项在检查时看似无害的技能,可能仅在遇到特定环境状态、资源或交互历史后才会暴露有害行为。现有扫描器主要依赖静态分析、预定义规则或单次语义判断,难以触发和归因此类条件性行为。我们提出SkillSentry,一种基于自适应蜜罐世界的动态安全测试框架。SkillSentry推断技能的预期能力边界,构建带有受控诱饵资源的LLM模拟环境,并自适应生成任务以探索其行为状态。随后,它将启用技能的轨迹与匹配的无技能执行进行比较,将可疑行为归因于源代码和验证过的执行轨迹,再做出最终决策。我们针对7种扫描器配置对SkillSentry进行评估,SkillSentry在标准基准测试上达到99.50%的召回率和96.26%的平均F1值;在语义保持规避场景下,其平均F1值达到92.95%,而最强基线的平均F1值为80.07%。我们的代码可在https URL获取。
英文摘要
External skills extend the capabilities of large language model agents, but also introduce an execution-time attack surface: a skill that appears benign under inspection may reveal harmful behavior only after particular environmental states, resources, or interaction histories are encountered. Existing scanners primarily rely on static analysis, predefined rules, or one-shot semantic judgments, making such conditional behavior difficult to elicit and attribute. We present SkillSentry, a dynamic safety-testing framework based on adaptive honey worlds. SkillSentry infers the intended capability boundary of a skill, constructs an LLM-simulated environment with controlled decoy resources, and adaptively generates tasks to explore its behavioral states. It then compares skill-enabled trajectories with matched no-skill executions, grounding suspicious behaviors in source code and verified execution traces before making a final decision. We evaluate SkillSentry against seven scanner configurations. SkillSentry achieves 99.50% Recall and 96.26% average F1 on standard benchmarks. Under semantics-preserving evasion, it reaches 92.95% average F1, compared with 80.07% for the strongest baselines. Our code is available at https://github.com/nizhangli062-jpg/SkillSentry-Adaptive-Honey-Worlds-for-Dynamic-Safety-Testing-of-Agent-Skills.