发表机构
The Hong Kong University of Science and Technology; South China Normal University(香港科技大学; 华南师范大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出测试时原型适配(TPA),一种无训练即插即用模块,可与多种开放词汇语义分割宿主结合,仅用少量未标注图像构建原型库,无需调优或更新参数即可提升分割准确率。
AI 中文摘要
开放词汇语义分割(OVSS)在无需额外标注监督的情况下,将预训练的CLIP编码器重新用于密集预测。现有方法要么通过重新设计CLIP的内部注意力,要么通过注入辅助视觉基础模型(VFM)的特征来改善CLIP的空间行为,这两种方法都需要访问宿主的内部计算,并针对其特定的前向传播进行定制。在本研究中,我们提出了测试时原型适配(TPA),这是一种无训练的即插即用模块,在输出层运行,不修改宿主的前向传播和权重。通过利用轻量级转导适配阶段,TPA从宿主在一小部分未标注部署域图像上的自身输出预测中识别出可信的锚点补丁,并将其冻结的DINO特征聚合为每类原型;在推理时,针对该冻结库的单次余弦相似度查找会提供一个辅助分数,与宿主的对数线性融合。TPA与五种代表性的OVSS宿主(涵盖注意力重新设计和VFM注入设计)、三种CLIP骨干网络、八个基准以及多种内部VFM选择相结合。在单一超参数设置下,无需针对每个宿主进行调优或参数更新,TPA始终能提升分割准确率,且在大多数基准上,仅需约10%的未标注部署域图像就足以构建有效的特征库。
英文摘要
Open-vocabulary semantic segmentation (OVSS) repurposes a pretrained CLIP encoder for dense prediction without additional labeled supervision. Existing methods improve CLIP's spatial behavior either by redesigning its internal attention or by injecting features from auxiliary vision foundation models; both require access to the host's internal computation and are tailored to its specific forward pass. In this work, we propose Test-time Prototype Adaptation (TPA), a training-free plug-in that operates at the output level, leaving the host's forward pass and weights unmodified. By leveraging a lightweight transductive adaptation phase, TPA identifies confident anchor patches from the host's own output predictions on a small pool of unlabeled deployment-domain images, and aggregates their frozen DINO features into per-class prototypes; at inference, a single cosine similarity lookup against this frozen bank provides an auxiliary score fused linearly with the host's logits. TPA composes with five representative OVSS hosts spanning attention-redesign and VFM-injection designs, across three CLIP backbones, eight benchmarks, and multiple internal VFM choices. Under a single set of hyper-parameters and without per-host tuning or parameter updates, TPA consistently improves segmentation accuracy, with as few as approximately 10% of unlabeled deployment-domain images sufficing for effective bank construction on most benchmarks.
Comments17 pages, 12 figures, preprint