语言模型中的无意上下文泄漏
Inadvertent Context Leakage in Language Models
浏览论文内容
中文总结 AI 辅助
该研究发现语言模型的上下文窗口会无意泄漏敏感用户信息,提出自适应攻击方法,在8个专有模型上验证了2、4位数秘密的重建准确率,还展示了两种实际攻击方式,指出泄漏是模型能力的副产品。
中文摘要 AI 辅助
为了让AI智能体超越简单的聊天功能发挥作用,它们必须存储敏感的用户上下文,如日历、凭证、健康记录和财务数据。我们研究模型上下文窗口中仅存在此类秘密是否会在模型的良性输出中引入隐藏的相关性,从而即使模型直接拒绝提取时也能进行重建。我们进一步研究对手是否可以主动设计提示来放大这种效果,利用模型作为隐蔽载体,通过看似无害的文本传输秘密。在这两种情况下,这种有限的泄漏都使用一种新颖的自适应攻击来利用,该攻击假设可以黑盒访问底层模型。在对八个专有模型进行的受控实验中,我们发现2位数的上下文秘密可以以近乎完美的准确率重建,4位数的秘密精确匹配率为82%,全部来自模型对普通、非对抗性请求产生的输出。我们观察到更强大的模型泄漏更多:更强的指令遵循能力会放大对上下文秘密的敏感性,这表明泄漏是能力的副产品,而非可修补的错误。我们展示这种泄漏可实现两种实际攻击:(1)训练后的分类器从常规自然语言输出中推断关于用户记忆的语义谓词(如健康状况、财务事件);(2)一个经RL训练的对手从生产式智能体中提取完整的社会安全号码。
英文摘要
For AI agents to be useful beyond simple chat, they must hold sensitive user context such as calendars, credentials, health records, and financial data. We study whether the mere presence of such secrets in a model's context window introduces hidden correlations into the model's benign outputs, allowing reconstruction even when the model correctly refuses direct extraction. We further study whether an adversary can actively engineer prompts that amplify this effect, using the model as a covert carrier to transmit secrets through seemingly innocuous text. In both cases, this limited leakage is exploited using a novel adaptive attack that assumes black-box access to the underlying model. In controlled experiments across eight proprietary models, we find that 2-digit in-context secrets are reconstructed with near-perfect accuracy and 4-digit secrets at 82\% exact match, all from outputs the model produces in response to ordinary, non-adversarial requests. We observe that more capable models leak more: stronger instruction-following amplifies sensitivity to in-context secrets, suggesting leakage is a byproduct of capability as opposed to a patchable bug. We show this leakage enables two practical attacks: (1) a trained classifier that infers semantic predicates about user memories (e.g., health conditions, financial events) from routine natural-language outputs, and (2) an RL-trained adversary that extracts full Social Security Numbers from a production-style agent.
发表机构
- FAIR, Meta Superintelligence Labs(Meta超级智能实验室FAIR)
- University of California, Berkeley(加利福尼亚大学伯克利分校)
- Google DeepMind(谷歌DeepMind)
机构由 AI 辅助整理,请以论文原文为准。