发表机构
The University of Western Australia(西澳大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对长时程LLM智能体,提出基于来源感知的执行图定义影响距离,揭示步数隐藏的安全暴露,支持运行时干预。
AI 中文摘要
长时程LLM智能体与不可信内容、持久记忆、外部状态和敏感工具交互。现有分析通常通过恶意输入与下游动作之间的执行步数来刻画攻击。我们表明,在有状态智能体中,时间上的遥远性可能高估安全隔离。我们引入一个基于来源感知的执行图,通过确定性状态、标识符和工具来源连接智能体事件,并定义影响距离$\DI$为从不可信源到敏感动作的最短结构路径。我们将其与序列距离$\DT$(有序轨迹中的最短注入-汇路径)进行比较。由于影响图包含每条序列边,$\DI \leq \DT$;$\Gap=\DT-\DI$衡量步数所隐藏的隔离。在来自OpenAI的\texttt{gpt-4o-mini}和\texttt{gpt-4o}以及Claude的Haiku 4.5和Sonnet 4.6的360条长时程AgentDojo轨迹中的454个注入-汇对中,96.9%的对$\Gap>0$,中位差距为9跳;在移除最大的仅来源边类别后,91.0%仍然解耦。在AgentDojo的银行套件中,来自377条轨迹的231对中的33.8%通过不同的来源机制解耦。在274个OpenAI对中,在控制$\DT$、攻击家族和后端后,$\Gap$不能独立预测攻击成功($\beta_{\Gap}=0.066$,$p=.088$)。在匹配阈值$k=2,3$时,基于$\DI$的确定性执行前门阻止了仅序列门遗漏的五个攻击汇,且没有额外的良性阻止,尽管配对增益不显著($p=.0625$)。因此,执行结构揭示了步数所隐藏的接近性,并可以支持有针对性的运行时干预。我们衡量候选影响路径而非因果归因。
英文摘要
Long-horizon LLM agents interact with untrusted content, persistent memory, external state, and sensitive tools. Existing analyses often characterize attacks by the number of execution steps between malicious input and a downstream action. We show that temporal remoteness can overstate security separation in stateful agents. We introduce a provenance-aware execution graph linking agent events through deterministic state, identifier, and tool provenance, and define \emph{influence distance} $\DI$ as the shortest structural path from an untrusted source to a sensitive action. We compare it with \emph{sequence distance} $\DT$, the shortest injection--sink path in the ordered trajectory. Since the influence graph contains every sequence edge, $\DI \leq \DT$; $\Gap=\DT-\DI$ measures the separation hidden by step count. Across 454 injection--sink pairs from 360 long-horizon AgentDojo trajectories over OpenAI's \texttt{gpt-4o-mini} and \texttt{gpt-4o} and Claude's Haiku 4.5 and Sonnet 4.6, $\Gap>0$ for 96.9% of pairs, with a median gap of 9 hops; 91.0% remain decoupled after removing the largest provenance-only edge class. On AgentDojo's banking suite, 33.8% of 231 pairs from 377 trajectories decouple through different provenance mechanisms. Among 274 OpenAI pairs, $\Gap$ does not independently predict attack success after controlling for $\DT$, attack family, and backend ($β_{\Gap}=0.066$, $p=.088$). At matched thresholds $k=2,3$, a deterministic $\DI$-based pre-execution gate blocks five attack sinks missed by a sequence-only gate with no additional benign blocking, although the paired gain is not significant ($p=.0625$). Execution structure therefore reveals proximity hidden by step count and can support targeted runtime intervention. We measure candidate influence pathways rather than causal attribution.