arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.20211cs.CRcs.MA

沉默即认可:LLM智能体流水线中的验证状态清洗

Silence Is Endorsement: Verification-Status Laundering in LLM Agent Pipelines

  • Illinois Institute of Technology(伊利诺伊理工学院)

机构由 AI 辅助整理,请以论文原文为准。

Yibo Hu

AI总结:

研究发现LLM智能体流水线中交接信息会丢失验证状态,导致风险动作批准率大幅上升,提出将授权来源作为结构化状态随声明传递以解决此问题。

AI中文摘要:

LLM智能体系统中的安全监控器通常根据摘要或存储的交接信息来判断动作,而非依据原始证据。这造成了一种简单但危险的故障模式:交接信息保留了动作已获授权的声明,却丢失了该声明从未被验证的事实。我们将此称为验证状态清洗。在九个开源权重监控器和两个托管模型上,我们保持动作和授权命题不变,仅移除声明周围未经核实的来源框架。这一改动使Llama-3.1-8B上风险动作的批准率从5%升至60%,在Qwen2.5-14B上从9%升至98%,两个托管模型也出现了同样显著的提升。该故障也出现在普通智能体流水线中。摘要器经常削弱验证状态,记忆压缩器常将其移除,而完整的提议者-摘要器-记忆-监控器流水线在三个下游监控器上将风险动作批准率提升至57%-81%。在WildGuard和ATBench上的实验显示,对于独立编写的有害和不安全请求,同样的模式依然存在:未经支持的授权声明使批准的可能性显著增加。明确指示监控器拒绝未经验证的授权并非可靠的跨模型修复方案:一些模型仍然易受攻击,而另一些则拒绝合法请求。因此,智能体系统应将授权来源作为结构化状态附加到声明上,并在整个流水线中携带该状态。

英文摘要:

Safety monitors in LLM agent systems often judge actions from summaries or stored handoffs, not from the original evidence. This creates a simple but dangerous failure mode: the handoff preserves the claim that an action is authorized while losing the fact that the claim was never verified. We call this verification-status laundering. Across nine open-weight monitors and two hosted models, the action and authorization proposition remain fixed while we remove the unverified provenance framing around the claim. This change raises approval for risky actions from $5\%$ to $60\%$ on Llama-3.1-8B and from $9\%$ to $98\%$ on Qwen2.5-14B, with similarly large shifts on both hosted models. The failure also emerges in ordinary agent pipelines. Summarizers frequently weaken the status, memory compressors often remove it, and a full proposer--summarizer--memory--monitor pipeline raises risky approval to $57$--$81\%$ across three downstream monitors. Experiments on WildGuard and ATBench show the same pattern on independently authored harmful and unsafe requests: unsupported authorization claims make approval substantially more likely. Explicitly instructing monitors to reject unverified authorization is not a reliable cross-model fix: some models remain vulnerable, while others reject legitimate requests. Agent systems should therefore carry authorization provenance as structured state attached to the claim throughout the pipeline.

↑