arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.28502cs.CR

无执行的识别:大语言模型智能体指令仲裁与外部控制中依赖配置的失效

Recognition Without Enforcement: Configuration-Dependent Failures in LLM Agent Instruction Arbitration and External Control

Jun Wen Leong

AI总结:

本研究发现大语言模型智能体存在指令仲裁的识别-执行缺口,提出外部参考监控器实现安全控制,证明安全智能体需外部执行而非仅依赖识别。

AI中文摘要:

大语言模型智能体在系统提示、用户、记忆及工具的各类指令间进行仲裁,但该仲裁不能被假定为能强制信任边界。我们识别出一种识别-执行缺口:源格式特征(角色模板位置、通道元数据、格式提示)可从模型激活中线性解码,且模型在被提示时能明确识别伪造的权限,然而某些配置仍会产生冲突的工具调用。我们在此特定的可解码源格式加口头检测意义上使用“识别”;交叉探测控制显示其并非统一的抽象信任表示。该缺口并非模型权重的固有属性。限制性策略和多样化提示可在相同模型上消除执行,而宽松配置及特定提示-模型对会产生确定性失效。在一次 fleet 评估(权限伪造:来自6个供应商的46个模型端点,包括开放权重模型;记忆冲突:48个模型)中,多样化新型攻击下的平均执行率为1.21%[0.5-2.1%](基于29个模型的14294次伪造试验的模型聚类置信区间),但漏洞集中在可复现单元且随部署窗口变化(每个指纹的窗口内范围达47个百分点)。提示层防御同样无法跨模型和自适应表述泛化。因此,我们将模型自我仲裁视为一种能力而非安全边界,并实现了外部参考监控器,其结合了经认证的源路由与能力门控工具执行。它确定性地拒绝所有测试过的伪造、篡改、重放及未签名请求,同时保留合法操作。一次独立的自适应红队测试发现了一个实现缺陷(一个已修复的时钟偏差准入),而非密码学旁路。安全智能体需要外部执行,而非仅更好的识别。

英文摘要:

LLM agents arbitrate among instructions from system prompts, users, memory, and tools, but this arbitration cannot be assumed to enforce trust boundaries. We identify a recognition-enforcement gap: source-format features (role-template position, channel metadata, formatting cues) are linearly decodable from model activations, and models can explicitly identify forged authority when prompted, yet some configurations still produce the conflicting tool call. We use "recognition" in this specific decodable-source-format-plus-verbalized-detection sense; crossed-probe controls show it is not a unified abstract trust representation. The gap is not an immutable property of model weights. Restrictive policies and diverse prompts can eliminate execution on the same models, while permissive configurations and particular prompt-model pairs yield deterministic failures. Across a fleet evaluation (authority spoofing: 46 model endpoints across 6 vendors including open-weight; memory conflict: 48 models), average execution under diverse novel attacks is 1.21% [0.5-2.1%] (model-clustered CI over 14,294 spoofed trials from 29 models), but vulnerability is concentrated in reproducible cells and shifts across deployment windows (up to 47pp within-window per-fingerprint range). Prompt-layer defenses likewise fail to generalize across models and adaptive formulations. We therefore treat model self-arbitration as a capability rather than a security boundary and implement an external reference monitor combining authenticated source routing with capability-gated tool execution. It deterministically rejects all tested forged, tampered, replayed, and unsigned requests while preserving legitimate operations. A separate adaptive red-team found one implementation flaw (a since-patched clock-skew admission), not a cryptographic bypass. Secure agents require external enforcement, not merely better recognition.

补充信息

↑