arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2412.04073cs.CV

TransAdapter:用于以特征为中心的无监督域适应的 Vision Transformer

TransAdapter: Vision Transformer for Feature-Centric Unsupervised Domain Adaptation

  • Ozyegin University(厄兹耶金大学)

机构由 AI 辅助整理,请以论文原文为准。

A. Enes Doruk, Erhan Oztop, Hasan F. Ates

更新

AI总结:

本文提出 TransAdapter,将 Swin Transformer 与图域判别器、自适应双重注意力和跨特征变换结合,用于以特征为中心的无监督域适应,并在多个基准上达到最先进性能。

AI中文摘要:

无监督域适应(UDA)旨在利用来自源域的有标签数据来解决无标签目标域中的任务,但这通常会受到显著域差距的阻碍。传统的基于 CNN 的方法难以充分捕捉复杂的域关系,这促使人们转向 Swin Transformer 等 vision transformer,后者在建模局部和全局依赖关系方面表现出色。在本文中,我们提出了一种利用 Swin Transformer 的新型 UDA 方法,该方法包含三个关键模块。Graph Domain Discriminator 通过图卷积和基于熵的注意力差异化来捕捉像素间相关性,从而增强域对齐。Adaptive Double Attention 模块将 Windows 与 Shifted Windows 注意力相结合,并通过动态重加权有效对齐长程和局部特征。最后,Cross-Feature Transform 修改 Swin Transformer 块以提升跨域泛化能力。大量基准测试证实了我们这种通用方法的最先进性能;该方法不需要任务特定的对齐模块,从而确立了其对多种应用的适应性。

英文摘要:

Unsupervised Domain Adaptation (UDA) aims to utilize labeled data from a source domain to solve tasks in an unlabeled target domain, often hindered by significant domain gaps. Traditional CNN-based methods struggle to fully capture complex domain relationships, motivating the shift to vision transformers like the Swin Transformer, which excel in modeling both local and global dependencies. In this work, we propose a novel UDA approach leveraging the Swin Transformer with three key modules. A Graph Domain Discriminator enhances domain alignment by capturing inter-pixel correlations through graph convolutions and entropy-based attention differentiation. An Adaptive Double Attention module combines Windows and Shifted Windows attention with dynamic reweighting to align long-range and local features effectively. Finally, a Cross-Feature Transform modifies Swin Transformer blocks to improve generalization across domains. Extensive benchmarks confirm the state-of-the-art performance of our versatile method, which requires no task-specific alignment modules, establishing its adaptability to diverse applications.

↑