发表机构
Meta; Max Planck Institute for Intelligent Systems; University of Tübingen; National University of Singapore(Meta; 马克斯·普朗克智能系统研究所; 图宾根大学; 新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究视觉和视觉语言模型中视觉特征隐私问题,提出TrustCLIP框架,以特征条件生成器为隐私对手,优化编码器特征与下游模块投影,降低生成式反演保真度并保持下游任务性能。
AI 中文摘要
视觉和视觉语言模型依赖的高级视觉表示存在隐私风险。本文通过重建视角重新审视该问题,提出TrustCLIP框架,将特征条件生成器视为隐私对手,学习编码器特征与下游模块间投影,优化以降低生成式攻击者的重建质量,同时保持下游任务信号,在多种场景验证了有效性。
英文摘要
Vision and vision-language models rely on high-level visual representations that are increasingly used across recognition, retrieval, and multimodal reasoning pipelines. However, recent advances in generative modeling have shown that such features can often be inverted, enabling realistic reconstructions of the underlying image and raising significant privacy risks. We revisit this problem through the lens of reconstruction and propose TrustCLIP, a reconstruction-driven framework that treats a feature-conditioned generator as an explicit privacy adversary. TrustCLIP learns a projection between encoder features and downstream modules that is explicitly optimized to degrade the reconstructions produced by generative attackers while retaining the necessary signals for downstream tasks. Unlike prior defenses that rely on discriminative privacy metrics, TrustCLIP directly optimizes against a generative reconstruction attacker, targeting a threat not captured by standard evaluation protocols. We demonstrate its effectiveness in both conventional classification and multimodal large language model pipelines. Across these settings, TrustCLIP consistently reduces the fidelity of generative inversions while maintaining downstream task performance. Project page: https://atnikos.github.io/trustclip/
Commentshttps://atnikos.github.io/trustclip/ Update affiliations