arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32203cs.CVcs.LG

基于核的CLIP引导:利用视觉-语言模型偏好

Kernel-Based Steering of CLIP with Vision-Language Model Preferences

Sajjad Ghiasvand, Haniyeh Ehsani Oskouie, Sina Mansouri, Mahnoosh Alizadeh, Farzan Farnia, Ramtin Pedarsani

首次发表
浏览论文内容

中文总结 AI 辅助

提出ASK方法,利用VLM偏好通过核匹配和低秩适配器引导CLIP,无需教师嵌入,提升检索mAP至75.0并保持零样本能力。

中文摘要 AI 辅助

大型视觉-语言模型(VLM)能够判断视觉相似性,但其判断无法直接作为紧凑的图像嵌入用于高效比较。我们研究如何将这些偏好迁移到CLIP中,同时保留其图像-文本能力。我们提出ASK,一种基于核的引导方法,通过从引发的成对判断中学习,无需访问教师嵌入或收集新的人类相似性标注。ASK在小图像组内构建正半定目标核,并将视觉核匹配与图像-文本分布锚定相结合。低秩适配器联合更新视觉和文本编码器,同时将预测正则化至冻结的CLIP。适应后,检索使用CLIP图像嵌入和余弦相似度,无需调用VLM。在五个图像领域、四个CLIP骨干网络和六个评判者上的实验评估了教师一致性、检索和识别保持。对于ViT-B/16,在适应排除的类别上,平均检索mAP从53.8提升至75.0,而使用KL锚定的DINOv2目标为71.7。联合适应编码器的平均零样本准确率从61.8%提升至62.4%,在12个基准和五个适应领域上取平均。提示提供了额外能力:选择学生学习的视觉区分。跨四个数据集的人类标注评估支持这种标准特定的控制。

英文摘要

Large vision-language models (VLMs) can judge visual similarity, but their judgments are not directly available as compact image embeddings for efficient comparison. We study how to transfer these preferences into CLIP while retaining its image--text capabilities. We introduce ASK, a kernel-based steering method that learns from elicited pairwise judgments without accessing teacher embeddings or collecting new human similarity annotations. ASK constructs positive semidefinite target kernels within small image groups and combines visual kernel matching with an image--text distributional anchor. Low-rank adapters jointly update the visual and text encoders while regularizing predictions toward frozen CLIP. After adaptation, retrieval uses CLIP image embeddings and cosine similarity, with no VLM calls. Experiments across five image domains, four CLIP backbones, and six judges evaluate teacher agreement, retrieval, and recognition retention. For ViT-B/16, mean retrieval mAP on classes excluded from adaptation increases from 53.8 to 75.0, compared with 71.7 for DINOv2 targets with KL anchoring. Mean zero-shot accuracy with jointly adapted encoders increases from 61.8\% to 62.4\%, averaged over 12 benchmarks and the five adaptation domains. Prompting provides an additional capability: selecting which visual distinctions the student learns. Human-annotated evaluations across four datasets support this criterion-specific control.

↑