隐形捷径:视觉编码器为何了解你的相机
Invisible Shortcuts: Why Vision Encoders Know Your Camera
浏览论文内容
中文总结 AI 辅助
该研究发现视觉编码器会利用像素级隐形元数据痕迹形成捷径,元数据-语义关联越强模型敏感性越高,提出的缓解策略可降低敏感性并提升分布外泛化能力。
中文摘要 AI 辅助
深度视觉模型会利用捷径,依赖与监督信号相关的线索。此前的研究聚焦于可见偏差,例如物体-背景或纹理关联。我们发现了捷径学习的另一个来源:像素级嵌入的隐形元数据痕迹,例如图像处理和照片采集相关的元数据。我们假设,大规模语义监督(无论是通过分类标签(ImageNet)还是十亿级别的描述文本(LAION))在预训练期间自然会引发元数据-语义关联,导致模型将低级信号转化为预测特征。通过引入受控的元数据-语义关联,我们证明更强的关联会使模型对元数据痕迹的敏感性系统性更高,且在元数据分布偏移下性能下降更明显。我们进一步探索了预训练期间及之后应用的缓解策略,这些策略不仅降低了对目标元数据的敏感性,也降低了对未见过的元数据的敏感性,同时不牺牲下游任务的性能。元数据敏感性也有积极的一面:它部分解释了一些编码器强大的生成图像检测能力,而对其进行缓解可以提升分布外泛化能力。代码:this https URL
英文摘要
Deep vision models exploit shortcuts, relying on cues that correlate with supervision signals. Prior work has focused on visible biases, such as object-background or texture correlations. We identify a different source of shortcut learning: invisible metadata traces embedded at the pixel level, for metadata such as image processing and photo acquisition. We hypothesize that large-scale semantic supervision, whether through categorical labels (ImageNet) or billion-scale captions (LAION), naturally induces metadata-semantics correlations during pretraining, leading models to convert low-level signals into predictive features. By introducing controlled metadata-semantics correlations, we show that stronger ones produce systematically higher sensitivity to metadata traces and larger performance degradation under metadata distribution shifts. We further explore mitigation strategies applied during and after pretraining that reduce sensitivity not only to targeted metadata but also to unseen ones, without sacrificing performance on downstream tasks. Metadata sensitivity also has a positive side: it partly explains the strong generated-image detection ability of some encoders, while its mitigation can improve out-of-distribution generalization. Code: https://github.com/ryan-caesar-ramos/visual-encoder-traces
发表机构
- Czech Technical University in Prague(布拉格捷克理工大学)
- The University of Osaka(大阪大学)
机构由 AI 辅助整理,请以论文原文为准。