arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.11113cs.CV

DiscoVL:通过正交对抗正则化揭示视觉-语言模型的解耦跨模态表示学习

DiscoVL: Unveiling Disentangled C ross-Modal Representation Learning via Orthogonal Adversarial Regularization for V ision-Language Models

Mengping Dong, Jinbao Li, Fei Li

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对预训练视觉-语言模型适配新下游场景时泛化能力不足的问题,提出DiscoVL框架,通过多分支低秩残差对齐器与正交对抗正则化实现解耦跨模态表示学习,在15个基准上的相关任务中性能优于现有最先进方法。

中文摘要 AI 辅助

预训练视觉-语言模型在各类感知任务中表现出色,但在不牺牲泛化能力的前提下将其适配到新的下游场景仍非易事。现有的参数高效提示学习方法常产生不一致的表示,且未考虑语义分布偏移问题。本研究提出DiscoVL,一种解耦跨模态表示学习框架,将正交对抗正则化与结构化跨模态对齐相结合,用于视觉-语言模型。为解决跨模态交互不足的问题,DiscoVL设计了多分支低秩残差对齐器,将表示分解为子空间,并在每一层实现视觉流与文本流之间的双向跨模态反馈。此外,传统三元组约束会使特征过拟合到类别质心,因此本研究设计了对抗三元组损失的正交正则化,以防止质心坍塌并大幅提升泛化能力。在15个基准上的评估表明,DiscoVL在基类到新类泛化、跨数据集评估和少样本学习方面均优于现有最先进方法,实现了一致的性能提升。

英文摘要

Pre-trained vision-language models excel across varied perception tasks, but adapting them to novel downstream settings without sacrificing generalization remains non-trivial. Existing parameter-efficient prompt learning method often yields inconsistent representations and fails to account for semantic distribution shifts. In this work, we present DiscoVL, a disentangled cross-modal representation learning framework that couples orthogonal adversarial regularization with structured cross-modal alignment for vision-language models. To address the insufficient cross-modal interaction, our DiscoVL designs a multi-branch low-rank residual aligner that decomposes representations into subspaces and enables bidirectional cross-modal feedback between visual and textual streams at each layer. Furthermore, while conventional triplet constraints overfit features to class centroids, we design an orthogonal regularization for adversarial triplet loss, which prevents centroid collapse and substantially boosts generalization. Evaluations on 15 benchmarks demonstrate that DiscoVL delivers consistent improvements over state-of-the-art methods for base-to-novel generalization, cross-dataset evaluation, and few-shot learning

发表机构

  • Shandong Artificial Intelligence Institute(山东人工智能研究院)
  • Qilu University of Technology (Shandong Academy of Sciences)(齐鲁工业大学(山东省科学院))
  • University of Florida(佛罗里达大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑