arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PISA:一种用于测试时开放词汇目标检测的伪个体源域特征适应框架

PISA: A Pseudo-Individual Source-Domain Feature Adaptation Framework for Test-Time Open-Vocabulary Object Detection

Ziyan He, Xiongtai Yang, Tao Wang

arXiv 2608.14142首次发表:更新:

AI 中文总结

针对开放词汇目标检测测试时适应的缺陷,提出PISA框架,利用CIFE、FAM、BAA模块,无需源域数据,在损坏基准上实现最优性能,提升检测精度。

AI 中文摘要

开放词汇目标检测测试时适应(OVOD-TTA)旨在解决预训练基础模型在遇到图像域偏移时出现的性能下降问题。现有的无数据源OVOD-TTA方法要么依赖精细的测试时信息进行重新评分,要么依赖伪标签进行自训练,当初始预测效果较差时会导致显著的精度下降。同时,大多数传统源域估计方法会恢复适用于分类任务的抽象、稀疏表示,但无法捕获检测所需的密集、具体特征。为解决这些问题,我们提出PISA,一种可无缝集成到开放词汇视觉骨干网络的新型无数据源OVOD-TTA方法。该方法的核心组件包括:不变性特征提取器(CIFE)、特征对齐模块(FAM)和多尺度对齐框架(BAA)。为捕获适用于检测的特征,我们开发CIFE以利用CLIP视觉特征在损坏图像上的不变性,确保对各种损坏的鲁棒性。我们进一步开发FAM和BAA,用于预训练和适应,将损坏不变性特征转换为接近原始源域特征的伪个体源域特征。通过这种方式,使用密集且具体的伪个体源域特征进行监督,而非不可靠的伪标签信号。在损坏的VOC-C、COCO-C和LVIS-C基准上,针对三个基础模型开展的实验表明,PISA显著提升了原始模型的定位精度和类别识别准确率。值得注意的是,PISA无需访问源域数据即可实现最先进的性能,在COCO-C上的AP@50%指标较现有方法超出3.92%。

英文摘要

Open-vocabulary object detection test-time adaptation (OVOD-TTA) aims to address the performance degradation that pre-trained base models suffer when encountering image-domain shifts. Existing source-free OVOD-TTA methods rely either on refined test-time information for re-scoring or on pseudo-labels for self-training, leading to significant accuracy degradation when initial predictions are poor. Meanwhile, most conventional source-domain estimation methods recover abstract, sparse representations suitable for the classification task, but fail to capture the dense, concrete features required for detection. To address these issues, we propose PISA, a novel source-free OVOD-TTA method that can be seamlessly integrated into open-vocabulary visual backbones. The core components of our method are the Corruption-Invariant Feature Extractor (CIFE), the Feature Alignment Module (FAM), and a multi-scale alignment framework (BAA). To capture detection-suitable features, we develop CIFE to exploit the invariance of CLIP's visual features across corrupted images, ensuring robustness against various corruptions. We further develop FAM and BAA for the pre-training and adaptation to transform the corruption-invariant features into pseudo-individual source-domain features that are close to the original source-domain features. In this way, dense and concrete pseudo-individual source-domain features are used for supervision instead of unreliable pseudo-label signals. Experiments on the corrupted VOC-C, COCO-C, and LVIS-C benchmarks across three base models demonstrate that PISA substantially improves both the localization precision and the category recognition accuracy of the original models. Notably, PISA achieves state-of-the-art performance without requiring access to source-domain data, surpassing existing methods by 3.92% in AP@50% on COCO-C.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑