发表机构
Zhejiang University; Tsinghua University; The University of Hong Kong; Hainan University; Anhui University; The Hong Kong Polytechnic University(浙江大学; 清华大学; 香港大学; 海南大学; 安徽大学; 香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
AgentSnare是一种轨迹自适应欺骗系统,通过动态构建诱饵环境,吸收渗透智能体工具调用、转移其轨迹并瓦解攻击,在CVE-Bench实验中成功阻止了所有真实目标被利用。
AI 中文摘要
大型语言模型(LLM)智能体通过观察-行动循环实现渗透测试自动化,基于工具返回的观察结果选择行动,这种依赖关系使防御者能够注入欺骗性观察结果来误导智能体的决策过程。然而,现有防御措施严重依赖攻击前预先植入环境的静态、孤立人工制品,高级智能体可逐步识别并绕过这些人工制品,最终将利用尝试重新聚焦于真实目标。为解决该问题,本文提出AgentSnare,一种轨迹自适应欺骗系统,可动态展开诱饵环境,持续引导渗透智能体远离真实目标。具体而言,AgentSnare采用人工制品构建策略模型,基于智能体的交互历史和诱饵状态构建候选人工制品,随后验证这些候选并将有效人工制品逐步纳入符合事实的诱饵环境,从而通过吸收智能体的工具调用延迟攻击,在诱饵内转移其进入后的轨迹,并通过诱导基于诱饵证据的完成报告来瓦解攻击。在15个CVE-Bench网络应用和3种攻击者模型的实验中,AgentSnare在诱饵中吸收了智能体46.8%的工具调用,将55.9%的进入后行动保留在诱饵中,同时90.0%的完成尝试基于诱饵证据;在所有45个攻击者-CVE配对中,在pass@3设置下无真实目标被成功利用。
英文摘要
Large language model (LLM) agents automate penetration testing through an observation-action loop, selecting actions based on observations returned by tools. This dependence allows defenders to inject deceptive observations that can mislead the agent's decision-making process. However, existing defenses rely heavily on static, isolated artifacts planted in the environment prior to an attack. Advanced agents can progressively recognize and bypass these artifacts, ultimately refocusing their exploitation attempts on the real target. To address this issue, we introduce AgentSnare, a trajectory-adaptive deception system that dynamically unfolds a decoy environment to continually steer the penetration agent away from the real target. Specifically, AgentSnare employs an artifact-construction policy model that constructs candidate artifacts conditioned on the agent's interaction history and decoy state. AgentSnare then validates these candidates and incrementally incorporates valid artifacts into a factually consistent decoy environment, thereby delaying the attack by absorbing its tool calls, diverting its post-entry trajectory within the decoy, and defusing it by inducing completion reports grounded in decoy evidence. Across 15 CVE-Bench web applications and three attacker models, AgentSnare absorbs 46.8% of the agent's tool calls in the decoy and retains 55.9% of post-entry actions there, while 90.0% of completion attempts are grounded in decoy evidence; across all 45 attacker-CVE pairs, no real target is successfully exploited at pass@3.