arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25711cs.CR

重构分布式风险:面向多轮智能体安全的轨迹条件动作生成

Reassembling Distributed Risk: Trajectory-Conditioned Action Generation for Multi-Turn Agent Safety

Yanbo Dai, Zhenlan Ji, Zongjie Li, Shuai Wang

首次发表
浏览论文内容

中文总结 AI 辅助

针对多轮分解攻击下使用工具的LLM智能体的安全风险,提出生成时防御方法ReDiR,通过轨迹级安全条件整合跨轮安全信息,将攻击成功率降至8%以下,可迁移且开销低。

中文摘要 AI 辅助

使用工具的大语言模型(LLM)智能体将安全风险从生成的文本扩展到影响外部系统的动作。在多轮分解攻击下,有害目标会被分散到多个看似合理的单独请求和工具调用中,仅从累积的轨迹中才能显现。现有防御要么依赖辅助在线推理来恢复长时序安全证据,要么在动作生成后进行评估,通常会产生额外的推理成本,或依赖特定运行时的动作表示。我们提出了Reassembling Distributed Risk(ReDiR,重构分布式风险),这是一种生成时防御方法,它将动作生成建立在轨迹级安全证据的条件上。在每个动作之前,ReDiR会将当前轨迹压缩为紧凑的潜在安全表示,并将其注入到冻结的基础模型中。该表示通过同模型跨视图监督学习得到,其中显式任务视图的安全行为为从原始多轮轨迹中恢复分布式安全证据提供监督。这种设计使ReDiR能够在生成过程中直接整合跨轮安全信息,无需依赖单独的动作级安全模块。我们在两个智能体安全基准、三个模型家族和八个未见过的工具领域上对ReDiR进行评估。ReDiR将攻击成功率降低到8%以下,可迁移到未见过的工具领域,并以较低的计算开销保留了良性动作的保真度。

英文摘要

Tool-using LLM agents extend security risks beyond generated text to actions that affect external systems. Under multi-turn decomposition attacks, a harmful objective can be distributed across individually plausible requests and tool calls, becoming apparent only from the accumulated trajectory. Existing defenses either rely on auxiliary online reasoning to recover long-horizon security evidence or assess actions after generation, often incurring additional inference cost or depending on runtime-specific action representations. We propose \emph{Reassembling Distributed Risk} (ReDiR), a generation-time defense that conditions action generation on trajectory-level security evidence. Before each action, ReDiR compresses the current trajectory into a compact latent safety representation and injects it into the frozen base model. The representation is learned through same-model, cross-view supervision, where safe behavior from an explicit task view provides supervision for recovering distributed safety evidence from the original multi-turn trajectory. This design enables ReDiR to integrate cross-turn security information directly within the generation process without relying on a separate action-level safety module. We evaluate ReDiR on two agent-safety benchmarks across three model families and eight held-out tool domains. ReDiR reduces attack success rates to below 8\%, transfers to unseen tool domains, and preserves benign fidelity with low computational overhead.

↑