发表机构
Corners Co., Ltd.(Corners Co., Ltd.)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
NOMOS是一款四阶段编译器,可将自然语言策略转为确定性工具调用网关,通过静态验证修复或拒绝部分候选规则,在τ²-bench等基准上大幅降低LLM智能体的工具调用违规率,且决策耗时极短,本地部署即可运行。
AI 中文摘要
使用工具的LLM智能体经常违反其部署时需执行的策略,且往往是隐性违反。现有防御方法包括手动编写规则、对每个动作查询LLM验证器,或通过重量级形式化机制编译策略。朴素编译存在缺陷:提取的规则会阻止满足自身前置条件的工具,或读取工具不具备的参数。NOMOS是一种四阶段编译器,可将自然语言策略转换为确定性工具调用网关;仅通过工具模式级别的静态验证(无需证明器、求解器或LLM)即可修复或拒绝37%(航空领域)和13%(零售领域)的候选规则,若不进行此步骤,多数已部署规则将无法运行。在未受防御的对话记录上重放编译后的规则,会标记出拒绝合法工作的绑定(某开发绑定拒绝了95.9%通过任务的调用),但未标记任何评估绑定。在τ²-bench上,该网关将状态变更调用中违反参考编码子句的比例从66.3%降至2.6%(航空领域)、从30.8%降至6.9%(零售领域),并在2≤k≤4时显著提升了航空领域的任务成功率;26B规模的本地部署编译结果与手动编写或前沿编译的规则相比无显著差异。与AgentDojo的已部署防御不同,它在银行领域达到了零攻击成功率(ASR),其中9个攻击家族可归为3条结构规则;在另外三个套件中,其ASR最高为3.6%,源于无工具调用可管控的目标以及某绑定允许了弱于对应子句的写入操作;第二个智能体模型Llama-3.3-70B在两个基准上均复现了该效果。决策过程无需LLM调用,耗时仅微秒级,仅存在与领域相关的良性效用成本;编译过程可在本地部署的开放权重模型gemma-4-26B上运行。
英文摘要
Tool-using LLM agents violate the policies they are deployed to enforce, often silently. Prior defenses hand-write rules, query an LLM verifier per action, or compile policies through heavyweight formal machinery. Naive compilation fails: extracted rules block the tool satisfying their own precondition, or read arguments their tool lacks. NOMOS, a four-pass compiler, turns a natural-language policy into a deterministic tool-call gate; static verification with tool-schema-level checks alone (no prover, solver, or LLM) repairs or rejects 37% (airline) and 13% (retail) of candidates, without which most shipped rules are inoperable. Replaying compiled rules over undefended transcripts flags bindings that refuse legitimate work (a development binding refused 95.9% of task-passing calls); no evaluation binding is flagged. On $τ^2$-bench the gate cuts violations of reference-encoded clauses among state-changing calls from 66.3% to 2.6% (airline) and 30.8% to 6.9% (retail), raising airline task success significantly for $2 \le k \le 4$; a 26B on-premise compilation is not significantly worse than hand-written or frontier-compiled rules. Unlike AgentDojo's shipped defenses, it reaches a zero attack success rate (ASR) on banking, where nine attack families collapse onto three structural rules. On the other three suites its ASR is at most 3.6%, from goals with no tool call to govern and one write admitted by a binding weaker than its clause; a second agent model, Llama-3.3-70B, reproduces the effect on both benchmarks. Decisions take microseconds without an LLM call, at a domain-dependent benign-utility cost; compilation runs on-premise on open-weight gemma-4-26B.
Comments28 pages, 5 figures, 19 tables. Submitted to IEEE Access