arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AgentTrap:针对自主渗透测试代理的状态性反馈欺骗

AgentTrap: Stateful Feedback Deception against Autonomous Penetration Testing Agents

Yuelin Wang, Jiongchi Yu, Yanbang Sun

arXiv 2610.02869首次发表:更新:

发表机构

Tianjin University; Nanyang Technological University(天津大学; 南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

AgentTrap提出首个针对自主渗透测试代理的闭环蜜罐,通过状态性欺骗和行为引导升级,将真实目标攻击成功率从95.8%降至79.2%,并有效诱出攻击者密钥。

AI 中文摘要

自主渗透测试代理通过持续调整其计划和行动来响应目标,从而执行多步攻击。作为一种常见的防御手段,蜜罐可以通过呈现诱饵服务来将此类代理从真实资产中转移,同时支持攻击追踪和主动反击。然而,传统蜜罐主要依赖静态工件和预定义响应,无法适应自主渗透测试代理不断演变的攻击策略。为此,我们提出了AgentTrap,这是首个专为自主渗透测试代理设计的闭环蜜罐。AgentTrap使用哨兵端点以避免良性干扰,基于受保护应用的状态性欺骗,以及行为引导的升级策略来维持交互,并在受控披露下收集代理侧行为证据。我们在一个部署的Web应用中对AgentTrap进行了评估,该应用包含一个真实应用端点和另一个配置了三种防御策略的独立蜜罐端点,测试对象为八个自主渗透测试代理。与无防御相比,AgentTrap将真实目标攻击总成功率从95.8%降至79.2%,并在18.8%的运行中成功诱出攻击者API密钥,优于静态欺骗和固定升级策略。此外,痕迹分析表明,对此类反击的抵抗能力同时依赖于模型层面对欺骗请求的识别以及架构层面敏感资源的隔离。

英文摘要

Autonomous penetration testing agents conduct multi-step attacks by continuously adapting their plans and actions to target responses. As a common defense, honeypots can be deployed to divert these agents from real assets by presenting decoy services, while also supporting attack tracing and active counterattacks. However, conventional honeypots rely primarily on static artifacts and predefined responses, leaving them unable to adapt to the evolving attack strategies of autonomous penetration testing agents. To this end, we present AgentTrap, the first closed-loop honeypot tailored for autonomous penetration testing agents. AgentTrap uses sentinel endpoints to avoid benign interference, stateful deception grounded in the protected application, and behavior-guided escalation to sustain engagement and collect agent-side behavioral evidence with controlled disclosures. We evaluate AgentTrap against eight autonomous penetration-testing agents in a deployed web application containing a real application endpoint and a separate honeypot endpoint configured under three defense strategies. Compared with no defense, AgentTrap reduces the aggregate real-target attack success rate from 95.8% to 79.2% and successfully elicits attacker API keys in 18.8% of the runs, outperforming static deception and fixed escalation. Furthermore, trace analysis shows that resistance to such counterattacks depends jointly on model-level recognition of deceptive requests and architecture-level isolation of sensitive resources.

Comments4 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑