发表机构
Huawei Technologies; EPITA(华为技术有限公司; 巴黎高等信息技术与应用科学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对AI智能体私钥易泄露问题,提出基于硬件密钥存储与五层零信任栈的架构,实验显示其攻击成功率降为0%且无误报。
AI 中文摘要
当前执行签名Git提交、认证API调用、颁发证书等密码学操作的AI智能体,将私钥存储在软件可访问的位置:明文文件、环境变量或容器内存中,任何拥有足够读取权限的进程都能提取原始密钥材料。近期一起生产事故凸显了其实际严重性:某广泛部署的框架通过邮件注入在不到五分钟内就泄露了私钥。我们的目标是对密钥使用实施密钥保密性和感知内容的授权。为此,我们用通过供应商中立的PKCS#11接口访问的硬件受限密钥取代驻留软件的密钥。硬件密钥存储(HSM、TPM、智能卡)在设备上执行密码学操作,主机仅通过不透明句柄接收结果。硬件隔离是主要贡献,它由围绕其的五层零信任实施栈实现,包括会话身份(SAGA)、范围边界(Smax)、语义验证(RAV)、污点跟踪和硬件执行边界。我们针对从AgentDojo的ImportantInstructionsAttack模板(Debenedetti等人,arXiv:2406.13352)衍生的12种注入场景进行评估,运行四个LLM模型:基线模式下三个模型受注入影响(gpt-oss-120b、Qwen2.5-72B、DeepSeek-V4-Flash,共n=192),基线攻击成功率(ASR)为19.3%[14.3%,25.4%];受保护的ASR为0%(Wilson 95%置信区间上限为2.0%),在四个良性任务场景中无误报。
英文摘要
AI agents increasingly sign Git commits, certify documents, and attest release artifacts on behalf of their operators, using private keys that live in software-accessible locations (plaintext files, environment variables, container memory) readable by any process the agent can reach. A widely deployed agent framework recently leaked its keys this way to a single email injection. Hardware keystores (HSM, TPM, smart card) keep the key on-device, but exposing the keystore as a tool an LLM agent can call moves the problem rather than removing it: once a signing session exists, the hardware cannot tell a request reflecting the operator's intent from one injected into content the agent read. We characterize this confused-deputy problem and build the five-layer Zero-Trust enforcement stack it requires, so that only requests consistent with the operator's committed intent reach the hardware. We evaluate on two attack planes. Prompt injection in content the agent reads (AgentDojo, three injection-following models, n=144) falls from an 18.1% baseline attack success rate to 0% under the full stack. Tool poisoning by a compromised MCP server (MCPTox) is contained identically: a hash comparison protects a pre-committed payload, and human-in-the-loop escalation contains autonomous requests with nothing pre-committed. A further probe delineates how far the semantic filter's protection extends: it detects a substitute document under an unrelated name, but an adversarially plausible substitute name defeats it in every trial we ran. We report this as a central finding: the architecture's guarantee never rests on the filter being right, only on a human being asked whenever nothing was committed in advance. The trade-off we characterize across both planes is that the less an operator can commit to in advance, the less deterministic the resulting guarantee, down to asking a human.
Commentsv2: substantially revised. Adds metadata-plane tool-poisoning evaluation (MCPTox), an adaptive substitution probe of the semantic filter, TPM 2.0 latency measurements, and per-layer ablations. Artifact link in the paper