发表机构
Tsinghua University; Nanyang Technological University; University of Washington(清华大学; 南洋理工大学; 华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对视觉-语言模型多标签测试时自适应的缓存方法PuRF,通过区域与缓存纯化解决现有方法的偏差与校准问题,在五项数据集上使ViT-B/32的mAP提升4.05%,性能优于现有最优方法。
AI 中文摘要
测试时自适应(TTA)已在单标签识别中得到广泛研究,可有效缓解分布偏移,尤其结合视觉-语言模型时效果显著。但现实图像常包含多个对象,更实用的多标签测试时自适应(MLTTA)至今受关注较少。近期基于缓存的TTA方法在效率和效果上表现良好,但直接扩展至多标签场景存在一对多映射问题:将关联共现对象的共享全局表示存储为类别级缓存原型,会引发主导标签偏差并损害缓存校准。引入区域级线索虽有助于分离类别特定证据,但这类区域线索在分布偏移下也可能不可靠,使其识别与利用颇具挑战。为解决这些问题,本文提出PuRF,一种新颖的纯化驱动型缓存方法,用于视觉-语言模型的多标签测试时自适应。具体而言,PuRF先执行区域纯化以识别可靠区域,为多标签识别提供全面区域线索并实现细粒度对齐;基于这些纯化区域,PuRF开展缓存纯化以增强缓存表示与适应性,其中 episodic 纯化构建判别性的基于区域的缓存, temporal 刷新进一步提升长期缓存适应性。实验表明,PuRF在五项数据集上针对ViT-B/32实现了4.05%的平均精度均值(mAP)提升,始终优于现有最优方法。
英文摘要
Test-time adaptation (TTA) has been widely explored in single-label recognition, effectively mitigating distribution shifts, especially when combined with vision-language models. However, real-world images often contain multiple objects, while the more practical multi-label test-time adaptation (MLTTA) has received little attention so far. Recent cache-based TTA methods have shown promising efficiency and effectiveness, yet directly extending them to multi-label scenarios suffers from a one-to-many mapping problem: a shared global representation entangling co-occurring objects is stored as class-wise cache prototypes, inducing dominant-label bias and compromised cache calibration. While introducing region-level cues helps isolate class-specific evidence, such regional evidence can also be unreliable under distribution shifts, making its identification and utilization non-trivial. To address these issues, we introduce PuRF, a novel PuRiFication-driven cache-based method for multi-label test-time adaptation of vision-language models. Specifically, PuRF first performs region purification to identify reliable regions, providing comprehensive regional cues for multi-label recognition and enabling fine-grained alignment. Based on these purified regions, PuRF conducts cache purification to enhance cache representation and adaptability, where episodic purification builds a discriminative region-based cache, and temporal refreshing further promotes long-term cache adaptability. Experiments demonstrate that PuRF consistently outperforms state-of-the-art methods, achieving a notable 4.05% mAP improvement on ViT-B/32 across five datasets.
CommentsAccepted by ECCV 2026