AI 中文总结
本文提出验证自主级别(VAL)分类法,比较VAL引导防御堆栈与主流直觉堆栈在LLM智能体安全中的效果,证明相同ASR零值可能源于结构保证或行为运气,零是结果而非保证。
AI 中文摘要
LLM智能体安全领域已涌现出密集的防御措施——提示加固、内容过滤、权限门控、沙箱——但没有任何框架能告诉部署者某项防御实际上保证了什么,或这种保证来自何处。我们将验证自主级别(VAL)——L0:LLM自我声明;L1:确定性规则;L2:客观真实值;L3/L4:可判定完备性;L5:不可能——应用于22种智能体安全防御措施;该分类法是可证伪的(在冻结卡片上10/10预测命中,已标记)。我们进行了首次受控的部署价值比较:在同等预算下,VAL引导的堆栈(确认门控+模式沙箱)与主流直觉堆栈(提示加固+关键词过滤)对比,涵盖50个场景、12种攻击变体、自适应/白盒/PAIR升级(约7,000次测试台调用;在AgentDojo/JADE上约10,000次测试框架调用)。VAL堆栈在1.000良性成功率下保持0.000攻击成功率(在AgentDojo银行任务上0.5% ASR对比未防御时的4.3%);直觉堆栈达到0.000 ASR但扼杀了所有良性动作——这是基于模型行为运气而非结构的安全。在攻击面难度递增的测试台上,直觉堆栈的零发生漂移(0->1.9%->6.2%,JADE上n=16),而VAL堆栈在其操作设计域(ODD)内保持稳定(0->0->0),其唯一突破是已披露的ODD外密码缺口(0.5%)。同一个零,两种不同的保证:零是结果,而非保证。
英文摘要
LLM-agent security has produced a dense landscape of defenses - prompt hardening, content filters, permission gates, sandboxes - yet no framework tells a deployer what a defense actually guarantees, or where that guarantee comes from. We apply Verification Autonomy Levels (VAL) - L0: LLM self-declaration; L1: deterministic rules; L2: objective ground truth; L3/L4: decidable completeness; L5: impossible - to 22 agent-security defenses; the taxonomy is falsifiable (10/10 prediction hits on frozen cards, flagged). We run the first controlled deployment-value comparison: at equal budget, a VAL-guided stack (confirmation gate + schema sandbox) versus a mainstream intuition stack (prompt hardening + keyword filter), 50 scenarios, 12 attack variants, adaptive/white-box/PAIR escalation (~7,000 testbed calls; ~10,000 harness calls on AgentDojo/JADE). The VAL stack holds 0.000 attack success at 1.000 benign success (0.5% ASR at 79.7% utility on AgentDojo banking vs 4.3% undefended); the intuition stack reaches 0.000 ASR but kills all benign actions - security by model-behavior luck, not structure. Across testbeds of rising attack-surface hardness the intuition stack's zero drifts (0->1.9%->6.2%, n=16 on JADE) while the VAL stack's holds within its ODD (0->0->0), its only breach a disclosed out-of-ODD password gap (0.5%). The same zero, two different guarantees: zero is an outcome, not a guarantee.
Commentsv1: applies Verification Autonomy Levels (VAL) to 22 agent-security defenses; first controlled deployment-value comparison (VAL-guided vs intuition stack) across three testbeds; ~17k LLM calls; an honest out-of-ODD boundary is reported. Writing was assisted by an AI language model; all experiments and research decisions are the author's own