AGATE:基于溯源的大语言模型智能体组合攻击运行时防御
AGATE: Provenance-Based Runtime Defense Against Compositional Attacks on LLM Agents
浏览论文内容
中文总结 AI 辅助
提出AGATE,一种基于授权和数据溯源门控的LLM智能体运行时防御系统,通过确定性检查阻止组合攻击,并在三个生产框架中验证了其可行性与局限性。
中文摘要 AI 辅助
大语言模型(LLM)智能体可能通过一系列普通操作产生有害后果。对此类行为进行判断,既需要确定授权其执行的权限来源,也需要追溯其携带数据的来源。我们提出AGATE,一种在经插桩的智能体框架边界处实施的授权与数据溯源门控机制。操作者声明和宿主批准事件为授权提供依据;委派操作受限于绑定到精确参数、具有有效期并限制使用次数的授权令牌。源注册将观察到的输入与后续传输关联起来,而效果账本则跟踪重复请求。确定性检查在决策路径中不依赖LLM即可做出判断,并保留其依据及执行证据以供取证重放。适配器集成了三个生产级框架——DeepSeek Harness、OpenCode和OpenClaw——无需修改宿主代码,将各宿主原生的观察点和否决点转换为统一的共享门控接口;判断核心在三个框架中完全相同,仅执行深度有所不同。我们的评估结合了153条实际演练的攻击链记录以及部署、效用和重建实验。部署观察揭示了工具声明和数据检查如何治理业务操作,包括通过参数重写实现的绕过。在11个良性文件处理场景中,有6个包含拒绝事件,揭示了基于内容溯源策略的效用成本。在63个净化场景上的252次运行中,重放与两个平台上所有63个场景的实时图投影一致。这些结果证明了基于溯源的运行时判断的可行性,并指出了内容转换、合法重用和观察覆盖范围是具体的局限。
英文摘要
LLM agents can produce harmful effects through sequences of ordinary operations. Judging such actions requires establishing both the authority that permits them and the origin of the data they carry. We present AGATE, an authorization and data-provenance gate at instrumented agent-harness boundaries. Operator declarations and host approval events ground authorization; delegated actions are constrained by grants that bind to exact parameters, expire, and permit a limited number of uses. Source registration connects observed inputs to subsequent transfers, while an effect ledger tracks repeated requests. Deterministic checks make decisions without an LLM in the decision path and retain their grounds with execution evidence for forensic replay. Adapters integrate three production harnesses -- DeepSeek Harness, OpenCode, and OpenClaw -- without modifying host code, translating each host's native observation and veto points into a single shared gate interface; the judgment core is identical in all three, and only enforcement depth differs. Our evaluation combines 153 exercised attack-chain records with deployment, utility, and reconstruction experiments. The deployment observations expose how tool declarations and data checks govern business actions, including a bypass through parameter rewriting. Six of eleven benign file-processing scenarios contain denial events, revealing the utility cost of content-based provenance policies. Across 252 runs on 63 sanitized scenarios, replay agrees with live graph projections for all 63 scenarios on each of two platforms. These results establish the feasibility of provenance-based runtime judgment and identify content transformation, legitimate reuse, and observation coverage as concrete limits.
发表机构
- Huazhong University of Science and Technology(华中科技大学)
机构由 AI 辅助整理,请以论文原文为准。