发表机构
University of Warwick; University of Hertfordshire(华威大学; 赫特福德大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM智能体结构化动作易受注入攻击的问题,提出一种认证多源完整性的安全认证器,通过有界破坏半径、最小命中集计数和确定性协调等机制,确保动作安全,并在实验验证中优于现有基线。
AI 中文摘要
LLM智能体越来越多地采取特权化的、往往不可逆的结构化动作,例如支付发票。它们从文档和工具输出中提取动作关键字段来组装每个动作,而对手可以篡改这些字段,且间接提示注入能驱动模型本身提取攻击者选择的值。当前的防御措施要么基于来源的信任标签进行门控,要么认证自由文本答案的质量。没有一种方法能在考虑共享上游来源的破坏预算下,认证耦合的、受策略约束的结构化动作的完整性。我们刻画了此类动作何时可安全认证,并给出了最大活性的安全认证器。它仅在每个字段通过其证据结构所支持的规则时才接受动作:对破坏不同的证据类别施加有界破坏半径,通过最小命中集计数,使得重新发布或清洗后的副本无法制造多数,对互补字段进行确定性协调,以及在证据使字段成为单一来源时提供可信锚点。我们形式化了两个鲁棒性概念,通过消融实验验证了每个机制,并测量了多源前提在制裁指定(70,966个实体)和软件供应链溯源(450个包)中成立的频率。在上界代理下,真正的相互印证在两者中都是少数现象,而朴素的证明计数会高估它,因为一旦计入共享来源,看似独立的见证者会坍缩为两个破坏不同的域。在真实智能体循环中的五个当前模型中,一次现实的注入欺骗了除一个外的所有模型,而一个朴素的智能体在大多数攻击下执行了欺诈性动作。认证器不接纳任何不安全动作,并在印证允许的地方恢复正确的值,而动作门控和溯源基线在我们的测试平台的每个世界中都被其空间中的某种攻击所破坏。
英文摘要
LLM agents increasingly take privileged, often irreversible structured actions, such as paying an invoice. They assemble each action from action-critical fields in documents and tool outputs that an adversary can corrupt, and indirect prompt injection can drive the model itself to extract attacker-chosen values. Current defenses gate on a source's trust label or certify free-text answer quality. None certifies the integrity of a coupled, policy-bound structured action under a corruption budget that accounts for shared upstream sources. We characterize when such an action is safely certifiable and give the maximally live safe certifier. It admits an action only when each field clears the rule its evidence structure supports: a bounded corruption radius over corruption-distinct evidence classes, counted by a minimum hitting set so that re-publishing or laundered copies cannot manufacture a quorum, deterministic reconciliation for complementary fields, and a trusted anchor where the evidence leaves a field single-sourced. We formalize two robustness notions, validate each mechanism by ablation, and measure how often the multi-source precondition holds on sanctions designations (70,966 entities) and software supply-chain provenance (450 packages). Under upper-bound proxies, genuine corroboration is a minority phenomenon in both, and naive attestation counting overstates it, since witnesses that look independent collapse to two corruption-distinct domains once shared origin is counted. Across five current models in a real agent loop, a realistic injection fools every model but one and a naive agent then executes the fraudulent action on most attacks. The certifier admits no unsafe action and recovers the correct value where corroboration permits, while action-gating and provenance baselines are broken in every world of our harness by some attack in its space.
CommentsAccepted to the UK AI Conference 2026 (UKAI 2026), to appear in Proceedings of Machine Learning Research. 14 pages, 3 figures, 7 tables. Code and data: https://github.com/anmolpandey299/certified-multisource-integrity
Journal refProceedings of the Fourth UK AI Conference 2026, Proceedings of Machine Learning Research 348 (2026) 134-147