arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27299cs.CRcs.SE

当上下文扎根:LLM 管控工具中的权限提升

When Context Gets Root: Privilege Escalation in LLM Harnesses

Xingbang He, Yuanwei Chen, Yi Qian, Haiyang Wei, Ligeng Chen, Zenan Fu, Linzhang Wang, Hao Wu, Bing Mao

首次发表
浏览论文内容

中文总结 AI 辅助

该研究发现LLM管控工具存在指令权限提升漏洞,攻击者可诱导智能体将低层级恶意内容提升至更高权限,通过多智能体机制在六个编码智能体管控工具上实现了全部13个涵盖多维度的攻击目标。

中文摘要 AI 辅助

指令层级是一种模型侧防御机制,它根据指令来源为其分配不同权限等级,这些等级用于限制哪些内容可引导模型行为。然而在智能体执行过程中,智能体管控工具会为每次模型调用构建上下文,这种构建可能会将低层级内容提升至更高指令层级,从而赋予其更大的模型侧权限。我们提出指令权限提升这一攻击类型:攻击者诱导智能体将低层级恶意内容提升至更高指令层级,被提升的内容会使智能体执行其在原始层级不会遵循的指令。我们通过多智能体机制在六个编码智能体管控工具上实现了13个攻击目标来评估该威胁,这些目标涵盖保密性、完整性、可用性和远程代码执行。在无限制动作执行条件下,攻击在全部六个管控工具上实现了全部13个目标;在自动权限审核条件下,攻击在提供该模式的全部三个管控工具上实现了全部13个目标。我们还利用管控工具提供的持久目标和定时任务复现了该漏洞,这些结果证明了指令权限提升的普遍性。

英文摘要

Instruction hierarchy is a model-side defense that assigns instructions different levels of privilege according to their sources. These levels constrain which content may direct model behavior. During agent execution, however, agent harnesses construct context for each model invocation. This construction can elevate low-level content to a higher instruction level and grant it greater model-facing privilege. We introduce instruction privilege escalation. In this attack, an attacker induces an agent to elevate low-level malicious content to a higher instruction level. The elevated content then causes the agent to execute instructions it would not follow at their original level. We evaluate this threat by using multi-agent mechanisms to achieve 13 attack objectives across six coding-agent harnesses. These objectives span confidentiality, integrity, availability, and remote code execution. With unrestricted action execution, the attacks achieve all 13 objectives on all six harnesses. Under automatic permission review, the attacks achieve all 13 objectives on all three harnesses that provide this mode. We further reproduce the vulnerability using harness-provided persistent goals and scheduled tasks. These results demonstrate the generality of instruction privilege escalation.

↑