发表机构
TU Darmstadt; Zuse School ELIZA(达姆施塔特工业大学; 祖塞ELIZA学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对统一多模态模型,提出隐私泄露水印(PLWs)攻击,使恶意模型提供者能利用对话历史在后续图像中隐藏可检测标记,实现隐私泄露,实验显示在1%假阳性率下可达100%真阳性率。
AI 中文摘要
多模态模型正日益转向统一架构,在共享的对话上下文中理解和生成文本、图像及其他模态。这种设计实现了跨模态的流畅交互,但也改变了隐私威胁模型:对话中某一部分泄露的信息,在模型随后于另一模态生成内容时可能仍然可访问。在用户依赖本地部署模型以保护隐私、并假设敏感交互仅局限于其设备的场景中,这一风险尤为令人担忧。我们引入了隐私泄露水印(PLWs):一种不可见、由触发器决定的水印,恶意模型提供者可以使其依赖于先前的聊天历史。借助这种对抗性干预,通常的隔离被打破:对话中早先提到的敏感关键词或语义线索,可导致后续无关图像携带隐藏但可检测的水印。PLWs对统一多模态模型的用户构成了一种新型威胁:被投毒的模型可以在保持实用性的同时,隐蔽地将图像生成转变为隐私泄露的渠道,即使部署在本地也是如此。在13个敏感属性触发器和两个模型家族中,PLWs在1%假阳性率下达到了高达100.0%的真阳性率。例如,在所有测试的对话间隔中,OmniGen2能检测到每一次先前披露的抑郁信息,同时仅将1%未含此类披露的图像错误标记。
英文摘要
Multimodal models are increasingly shifting toward unified architectures that understand and generate text, images, and other modalities within a shared conversational context. This design enables fluid interaction across modalities, but it also changes the privacy threat model: Information revealed in one part of a conversation may remain accessible when the model later generates content in another modality. This risk is particularly concerning in settings where users rely on locally deployed models for privacy, assuming that sensitive interactions remain confined to their device. We introduce Privacy-Leaking Watermarks (PLWs): invisible, trigger-dependent watermarks that a malicious model provider can condition on prior chat history. With this adversarial intervention, the usual separation breaks: a sensitive keyword or semantic cue mentioned earlier in the conversation can cause a later, unrelated image to carry a hidden yet detectable watermark. PLWs pose a novel threat to users of unified multimodal models: A poisoned model can retain utility while covertly turning image generation into a channel for privacy leakage, even when deployed locally. Across 13 sensitive-attribute triggers and two model families, PLWs reach up to 100.0% TPR at 1% FPR. For example, across all tested conversational separations, OmniGen2 detects every prior disclosure of depression while falsely flagging only 1% of images generated without such a disclosure.
CommentsCode: https://github.com/multimodal-ai-lab/PLW