arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

无需越狱的上下文推断攻击

Context Inference Attacks Without Jailbreaks

Prince Jha, Samuele Poppi, Nils Lukas

arXiv 2609.01663首次发表:更新:

发表机构

Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(穆罕默德·本·扎耶德人工智能大学(MBZUAI))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对智能体AI系统,提出无需越狱的上下文推断攻击,在三种场景下均能实现较高的上下文泄露率,刻画了泄露与相关参数的关系,揭示了智能体在隐私保护控制下仍存在的风险。

AI 中文摘要

智能体AI系统正越来越多地被部署用于在推理阶段处理敏感数据,比如医疗记录或金融文档,这些数据会在系统回答前被组装成隐藏的“上下文”。现有研究主要通过“越狱”攻击来考察隐私风险,这类攻击诱导模型直接披露敏感内容,但很大程度上忽略了智能体场景——即上下文是由智能体自身的工具调用组装而成的。我们的研究显示,尽管我们测试的智能体受到了相关控制(包括不披露上下文的指令、对数抑制、上下文稀释),但它们仍然容易遭受隐藏上下文的泄露。例如,一个执行良性用户查询的网页浏览智能体,仍然会携带可被利用的、关于其上下文中静默加载的记录的信号。我们通过安全博弈引入并形式化了“上下文推断攻击”,并在攻击者知识逐步减少、上下文交付方式日益间接的三种场景下对其进行评估:已知上下文、未知上下文、智能体通过自身工具调用检索到的上下文。我们区分了灰盒场景(在此场景中,目标模型用于对观测结果打分)和黑盒场景(在此场景中,攻击者用其控制的代理模型打分)。我们进一步刻画了泄露情况如何随查询预算、上下文大小和目标模型大小变化。单次攻击无需修改即可贯穿所有三种场景:针对已知上下文,在小型候选集上达到100%的ASR,在1024个候选集上达到63%;当模板和周围记录未知时,达到78.9的AUROC;当14B代理模型对32B目标模型打分时,达到92.5的AUROC;当记录作为智能体的检索返回结果时,达到81.8的AUROC,而对应的随机概率分别为1/|Z|和50。

英文摘要

Agentic AI systems are increasingly deployed to process sensitive data at inference time, such as healthcare records or financial documents assembled into a hidden \emph{context} before the system answers. Prior work has studied privacy risks primarily through \emph{jailbreaking} attacks that induce models to directly disclose sensitive content, but has largely overlooked the agentic setting where the context is assembled by the agent's own tool calls. We show that the agents we evaluate remain vulnerable to hidden-context leakage despite the controls we test against them, namely an instruction not to disclose the context, logit suppression, and context dilution. For instance, a web-browsing agent answering benign user queries still carries exploitable signals about records silently loaded into its context. We introduce and formalize \emph{context-inference attacks} through a security game and evaluate three settings under decreasing attacker knowledge and increasingly indirect delivery of the context: a known context, an unknown context, and a context the agent retrieves through its own tool calls. We distinguish a grey-box setting, in which the target model is used to score observations, from black-box settings in which the attacker scores with a surrogate it controls. We further characterize how leakage varies with query budget, context size, and target-model size. A single attack carries through all three settings without modification, reaching $100\%$ ASR on small candidate sets and $63\%$ at $1024$ candidates against a known context, $78.9$ AUROC when the template and surrounding records are unknown, $92.5$ AUROC when a 14B surrogate scores a 32B target, and $81.8$ AUROC when the records arrive as an agent's retrieval returns, against chance rates of $1/|\mathcal{Z}|$ and $50$ respectively.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑