发表机构
Fudan University; East China University of Science and Technology; Shanghai University of Electric Power(复旦大学; 华东理工大学; 上海电力大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多模态模型从照片推断地理位置导致的隐私泄露,提出基于扩散模型潜在空间扰动与GeoCLIP代理的主动防御方法,实现强黑盒迁移性与图像质量保持。
AI 中文摘要
多模态大型推理模型(MLRMs)在复杂视觉理解方面展现出卓越能力。然而,这种强大能力也引入了一个关键但尚未充分探索的隐私威胁:攻击者可以利用MLRMs,通过对建筑风格、植被和光照条件等微妙视觉线索进行结构化推理,从用户随意分享的照片中精确推断其地理位置。在本工作中,我们对MLRM驱动的地理定位隐私泄露进行了系统性研究。我们首先揭示,基于拒绝的防护措施严重不足,因为精心设计的越狱提示可将模型响应率提升至100%。我们进一步发现,现有的防御方法(向共享图像注入不可感知的扰动)受限于其像素空间优化的固有结构缺陷,导致黑盒迁移性下降并产生明显的视觉伪影。基于这些发现,我们提出了一种基于扩散的框架,提供针对地理定位隐私泄露的定向、主动防御。通过在反向采样过程中向扩散模型的潜在空间注入扰动,我们的方法直接作用于高层语义表征,从而从构造上解决了有效性与实用性之间的瓶颈。我们进一步以GeoCLIP(一种与GPS坐标显式对齐的模型)作为代理来锚定优化过程,以精确定位并破坏MLRMs用于位置推断的地理信号。这种定向语义破坏产生了显著更强的黑盒迁移性,同时保持了感知图像质量,为社交媒体平台提供了无缝集成方案。
英文摘要
Multimodal large reasoning models (MLRMs) have demonstrated remarkable capabilities in complex visual understanding. However, this very power introduces a critical yet underexplored privacy threat: adversaries can exploit MLRMs to precisely infer users' geographic locations from casually shared photographs, by performing structured reasoning over subtle visual cues such as architectural styles, vegetation, and lighting conditions. In this work, we present a systematic study of MLRM-driven geolocation privacy leakage. We first reveal that refusal-based safeguards are critically insufficient, as carefully crafted jailbreak prompts can raise model response rates to 100%. We further identify that existing defenses, which inject imperceptible perturbations into shared images, suffer from structural limitations intrinsic to their pixel-space optimization, resulting in degraded black-box transferability and pronounced visual artifacts. Motivated by these findings, we propose a diffusion-based framework that provides targeted, proactive defense against geolocation privacy leakage. By injecting perturbations into the latent space of a diffusion model during reverse sampling, our method operates directly on high-level semantic representations, thereby resolving the effectiveness-utility bottlenecks by construction. We further ground our optimization with GeoCLIP, a model explicitly aligned with GPS coordinates, as a surrogate to pinpoint and disrupt the geographic signals that MLRMs exploit for location inference. This targeted semantic disruption yields significantly stronger black-box transferability while preserving perceptual image quality, offering a seamless integration on social media platforms. Code is available at https://github.com/RachelWolowitz/Hiding_in_plain_sight.
CommentsNDSS 2027