arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越预定义汇点:面向LLM智能体的安全感知依赖分析

Beyond Predefined Sinks: Security-Aware Dependency Analysis for LLM Agents

Hang Cui

arXiv 2610.03014首次发表:更新:

AI 中文总结

针对LLM智能体安全分析中仅依赖预定义汇点不足的问题,提出AgentSecGraph框架,构建安全感知依赖图并引入AgentSecBench基准,实验证明该方法能有效区分安全敏感行为。

AI 中文摘要

基于大型语言模型(LLM)的智能体越来越多地将模型生成的决策与安全敏感的软件能力相连接,例如命令执行、文件系统访问、网络通信、浏览器控制和外部工具。现有的分析通常使用预定义的安全敏感操作作为锚点,但仅凭操作身份不足以确定安全影响。我们提出了AgentSecGraph,一个安全感知的静态分析框架,为每个安全敏感操作构建以候选者为中心的安全感知智能体依赖图(Security-ADG)。该框架将操作身份与智能体相关性、来源和依赖证据、信任边界上下文、防护证据以及外部效应语义相结合。我们进一步引入了AgentSecBench,一个包含67个真实世界LLM智能体仓库的语料库,涵盖11个生态系统和37,542个源文件。当前分析器在65个仓库中识别出23,866个静态安全敏感操作候选者,并为每个候选者生成一个Security-ADG工件。全语料库分析为9,821个候选者(41.15%)恢复了来源到操作的依赖证据,为3,075个候选者(12.88%)恢复了潜在防护证据,分析耗时50.8分钟。使用一个独立的基于复现的评估层,我们在13个仓库中确定了22个安全敏感行为:一个已确认的漏洞,一个待披露的候选者,以及20个受防护的行为。在九个保留案例中,Security-ADG保留了91.1%的参考上下文和所有五个观察到的防护,而仅汇点视图保留了20.0%,简化ADG保留了40.0%。这些结果表明,安全感知的依赖和上下文证据能够实现仅凭敏感操作身份无法恢复的区分。

英文摘要

Large language model (LLM)-based agents increasingly connect model-generated decisions to security-sensitive software capabilities such as command execution, filesystem access, network communication, browser control, and external tools. Existing analyses often use predefined sensitive operations as anchors, but operation identity alone is insufficient to determine security implications. We present AgentSecGraph, a security-aware static analysis framework that constructs a candidate-centered Security-Aware Agent Dependency Graph (Security-ADG) for each security-sensitive operation. It augments operation identity with agent relevance, source and dependency evidence, trust-boundary context, guard evidence, and external-effect semantics. We further introduce AgentSecBench, a corpus of 67 real-world LLM-agent repositories spanning 11 ecosystems and 37,542 source files. The current analyzer identifies 23,866 static security-sensitive operation candidates across 65 repositories and emits one Security-ADG artifact per candidate. Corpus-wide analysis recovers source-to-operation dependency evidence for 9,821 candidates (41.15%) and potential guard evidence for 3,075 (12.88%), completing in 50.8 minutes. Using a separate reproduction-backed evaluation layer, we establish 22 security-sensitive behaviors across 13 repositories: one confirmed vulnerability, one pending disclosure candidate, and 20 guarded behaviors. In nine held-out cases, Security-ADG preserves 91.1% of the reference context and all five observed guards, compared with 20.0% for a sink-only view and 40.0% for a simplified ADG. These results show that security-aware dependency and contextual evidence enable distinctions that cannot be recovered from sensitive-operation identity alone.

Comments29 pages, 3 figures, including supplementary appendices. Preprint

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑