arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39549cs.CRcs.AIcs.CL

投机性安全蜜罐:面向多轮智能体攻击的主动防御

Speculative Safety Honeypot: Toward Proactive Defense Against Multi-turn Agent Attacks

Zezhong Wang, Xueyang Tang, Rui Lian, Yang Lou, Heqing Huang

首次发表
浏览论文内容

中文总结 AI 辅助

针对多轮智能体攻击,提出投机性安全蜜罐(SSH)框架,利用多智能体模拟与推测-验证工作流提前预测风险,降低误报并提升防御韧性。

中文摘要 AI 辅助

随着大型语言模型(LLM)智能体日益部署于复杂环境中,多轮交互攻击已成为一项重大的安全挑战。现有检测方法通常依赖历史上下文。然而,这种回顾性逻辑难以识别被分散在多轮中、以隐藏未来风险的深层恶意意图。受投机性解码的启发,我们提出了投机性安全蜜罐(SSH)框架。SSH 采用由小型 LLM 组成的多智能体模拟系统,构建动作层面的“推测-验证”工作流。在推测阶段,SSH 预测目标智能体的未来行为,并异步构建轨迹树,以提前暴露潜在风险。在验证阶段,系统利用目标智能体的真实动作对轨迹树进行校准和剪枝,有效降低误报率。作为即插即用组件,SSH 为现有检测器提供了超越当前交互片段的丰富决策冗余。通过基于整个轨迹树的演化而非单一时间点进行风险判断,系统降低了对单个检测组件绝对精度的依赖。这提升了智能体系统针对复杂时间性攻击的防御韧性和预警提前量。

英文摘要

As Large Language Model (LLM) agents are increasingly deployed in complex environments, multi-turn interaction attacks have become a significant security challenge. Existing detection methods typically rely on historical context. However, this retrospective logic struggles to identify deep malicious intents that are split across turns to hide future risks. Inspired by speculative decoding, we propose the Speculative Safety Honeypot (SSH) framework. SSH uses a multi-agent simulation system composed of small LLMs to build an action-level speculate-and-verify workflow. In the speculation stage, SSH predicts future behaviors of the target agent and asynchronously builds a trajectory tree to expose potential risks in advance. In the verification stage, the system uses the target agent's real actions to calibrate and prune the trajectory tree, effectively reducing false positives. As a plug-and-playable component, SSH provides existing detectors with rich decision redundancy beyond the current interaction slice. By judging risk based on the evolution of the entire trajectory tree rather than a single point in time, the system reduces the reliance on the absolute precision of individual detection components. This improves the defense resilience and the warning lead-time of agent systems against complex temporal attacks.

发表机构

  • Huawei Technologies Co., Ltd(华为技术有限公司)

机构由 AI 辅助整理,请以论文原文为准。

↑