Bypassing Prompt Guards in Production with Controlled-Release Prompting
绕过生产环境中的提示守卫:受控释放提示攻击
机构 * UC Berkeley(加州大学伯克利分校) ; zkBricks Inc(zkBricks公司) ; Ethereum Foundation(以太坊基金会) ; NYU Shanghai(纽约大学上海分校)
专题命中 推理评测 :reasoning(abstract);分类 cs.LG
AI总结 针对AI对齐的提示过滤存在理论上的不可能性,本文提出受控释放提示攻击,利用轻量级输入过滤器与主模型之间的资源不对称性,在实际部署的大语言模型系统中成功绕过提示守卫。
Comments Accepted to USENIX Security 2026