arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

重新校准的对比损失用于视觉-语言模型中的变换感知提示条件化

Re-calibrated Contrastive Loss for Transformation-Aware Prompt Conditioning in Vision-Language Models

Seungmin Oh, Seunghun Kang, Jongbin Ryu

arXiv 2609.06967首次发表:更新:

发表机构

Ajou University(明知大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出变换感知提示条件化和重新校准对比损失,将同类样本视为正样本,缓解梯度稀释,提升视觉-语言模型在迁移、分布偏移和少样本场景下的性能。

AI 中文摘要

确保视觉-语言模型在不损害其泛化性能的情况下进行有效的迁移学习至关重要。然而,许多现有方法忽视了数据特征,简单地复用预训练期间采用的训练策略。具体来说,它们将同类样本视为不同的实例,并独立于配对的文本提示来变换图像,这使得模型学习更加困难。我们通过变换感知的提示条件化和重新校准的对比损失来解决这些局限性。固定的文本描述符识别应用于配对图像的变换,提供变换级别的一致性而不改变类别语义。这种设计在变换级别对齐图像和文本分支,使得表示更丰富,同时保持模型的泛化能力。此外,我们的损失函数在软目标交叉熵中,当每个锚点有多个有效正样本时,缓解了正梯度稀释问题。在迁移过程中,我们的方法将同类样本视为正样本而非不同实例,使模型能够更有效地学习领域特定特征。在分布偏移、迁移学习和少样本设置下的实验表明,我们的方法相对于现有方法具有一致的改进。我们方法的源代码可在该 https URL 获取。

英文摘要

Ensuring effective transfer learning for vision-language models without compromising their generalization performance is crucial. However, many existing methods overlook data characteristics and simply reuse the training strategies adopted during pre-training. Specifically, they treat same-class samples as distinct instances and transform images independently of their paired text prompts, which makes model learning more difficult. We address these limitations through transformation-aware prompt conditioning and a re-calibrated contrastive loss. Fixed text descriptors identify the transformations applied to paired images, providing transformation-level consistency without altering class semantics. This design aligns the image and text branches at the transformation level, enabling richer representations while preserving the models' ability to generalize. In addition, our loss function mitigates positive-gradient dilution in soft-target cross-entropy when each anchor has multiple valid positives. During transfer, our approach treats same-class samples as positives rather than distinct instances, enabling the model to learn domain-specific features more effectively. Experiments across distribution shift, transfer learning, and few-shot settings demonstrate consistent improvements over existing approaches. Source code for our method is available at https://github.com/SoongE/ReCalCon.

CommentsAccepted by British Machine Vision Conference (BMVC) 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑