arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

KALE:通过损失均衡实现网络规模下CLIP与DINOv2的稳定对齐的核对齐方法

KALE: Kernel Alignment with Loss Equilibration for Stable CLIP-DINOv2 Alignment at Web Scale

Michał Pawłowicz

arXiv 2607.18885首次发表:更新:

AI 中文总结

研究在网络规模数据上CLIP与DINOv2对齐问题,引入KALE损失均衡控制器,自适应调整对齐权重,无需数据集微调,在CC12M子集实验中提升了模型图像-文本检索及线性探测性能,零样本性能显著提高。

AI 中文摘要

基于核的CLIP与以视觉为中心的教师模型(如DINOv2)对齐(KUEA),在使用固定权衡权重在精心策划的ImageNet-1K上进行调整时,可改善CLIP的视觉表示,同时保持文本编码器的兼容性。研究发现这种方法在嘈杂的网络规模数据(CC12M)上效果不佳,对齐项的加权贡献降至清洁项的约0.2%。为此引入KALE,一种损失均衡控制器,它跟踪两种损失并自适应地将对齐权重重新调整到目标比率,无需针对每个数据集进行调整。通过实验表明,在CC12M子集上,对齐模型保留了图像-文本检索能力,并可重复性地提高了SVHN线性探测性能,零样本性能在标准11数据集平均水平上比CLIP提高了2.00,超过了KUEA的1.29。

英文摘要

Kernel-based alignment of CLIP toward a vision centric teacher such as DINOv2 (KUEA) improves CLIP's visual representations while preserving text-encoder compatibility, using a fixed trade-off weight tuned on curated ImageNet-1K. We ask whether this transfers to noisy, web-scale data (CC12M) and find that it does not: the alignment term's weighted contribution falls to about 0.2% of the clean term, so under any fixed weight its gradient is effectively inert. We introduce KALE, a loss-equilibration controller that tracks both losses and adaptively rescales the alignment weight toward a target ratio, restoring the signal with no per-dataset tuning; reaching balance requires increasing the weight by roughly four orders of magnitude, and the required value is configuration-dependent, so no fixed scalar suffices. We characterize the resulting regime: a bounded high learning rate and a decaying schedule with a moderate floor are needed for stability, and the controller equilibrates rather than diverging. On a 3.3M-image CC12M subset, the aligned model preserves image-text retrieval and reproducibly improves SVHN linear probing; zero-shot improves by +2.00 over CLIP on the standard 11-dataset average, exceeding KUEA's +1.29. We report all results with explicit run-to-run variance and base our conclusions on the metrics that are stable across runs.

Comments13 pages, 7 figures, 6 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑