arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CLIP 所知道但无法表达的:从冻结的中间特征中恢复否定信息

What CLIP Knows but Cannot Say: Recovering Negation from Frozen Intermediate Features

Chen-Yi Lu, Yueh-Shao Chen, Somali Chaterji

arXiv 2607.23271首次发表:更新:

发表机构

Purdue University(普渡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

对比视觉语言模型如CLIP对否定不敏感,研究发现是表征坍缩所致。提出轻量级事后校正系统PeakPatch,在不改变预训练权重下恢复否定信号,通过特定网络提取信号、预测偏差向量等,经实验验证其有效性及泛化性。

AI 中文摘要

对比视觉语言模型如 CLIP 将语义相反的短语(如“一只狗”与“不是一只狗”)映射到几乎相同的嵌入中,使其对否定不敏感。我们将此失败归因于一种称为表征坍缩的现象。通过跟踪 CLIP 文本编码器中的成分差异和视觉对齐,表明中间层构建成分句法,但最终层随着视觉对齐增加而坍缩此结构,产生句法盲的最终表示。为在不改变预训练权重的情况下恢复丢失的否定信号,我们提出 PeakPatch,一种轻量级事后校正系统。它在成分峰值处拦截编码器,通过嵌入校正网络(ECN)利用交叉注意力从峰值层提取特定于否定的信号,并预测偏差向量重新注入最终层嵌入空间,互补的分数校正网络(SCN)预测判别任务的有界标量分数偏移。两个模块端到端联合训练,CLIP 参数冻结,仅添加 520 万个参数(占主干的 3.5%)并保留标准余弦相似性接口。在 NegBench 上,PeakPatch 在 COCO MCQ 上达到 74.3%(比 CLIP 高 35.1,比最佳编码器微调方法高 17.8),在 VOC MCQ 上达到 65.5%,在完全分布外的否定检索上优于所有微调基线,校正后的嵌入还转移到文本到图像生成并在不同主干上泛化。

英文摘要

Contrastive vision-language models such as CLIP map semantically opposite phrases (e.g., "a dog" vs. "not a dog") to nearly identical embeddings, rendering them insensitive to negation. We attribute this failure to a phenomenon we call Representational Collapse: by tracking compositional divergence and visual alignment across the CLIP text encoder, we show that middle layers build compositional syntax, but the final layers collapse this structure as visual alignment rises, producing a syntax-blind final representation. To recover the lost negation signal without altering pretrained weights, we propose PeakPatch, a lightweight post-hoc correction system that intercepts the encoder at its compositional peak. An Embedding Correction Network (ECN) uses cross-attention to extract a negation-specific signal from the peak layer, anchored to a stable baseline, and predicts a deviation vector that re-injects the lost syntax into the final-layer embedding space. A complementary Score Correction Network (SCN) predicts bounded scalar score offsets for discriminative tasks. Both modules are trained jointly end-to-end while all CLIP parameters remain frozen, adding only 5.2M parameters (3.5% of the backbone) and preserving the standard cosine similarity interface. On NegBench, PeakPatch achieves 74.3% on COCO MCQ (+35.1 over CLIP, +17.8 over the best encoder fine-tuning method) and 65.5% on VOC MCQ, while outperforming all fine-tuning baselines on fully out-of-distribution negation retrieval despite training only 3.5% of the parameters. The corrected embeddings also transfer to text-to-image generation (+18.4 negation score) and generalize across ViT-B/32, ViT-L/14, and SigLIP backbones. Project URL: https://stevencylu.github.io/PeakPatch/.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑