arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向开放词汇语义分割中视觉-语言模型的持续测试时适应

Towards Continual Test-Time Adaptation of Vision-Language Models in Open-Vocabulary Semantic Segmentation

Chandler Timm C. Doloriel, Yunbei Zhang, Sarthak Kumar Maharana, Muhammad Salman Siddiqui, Tor Kristian Stevik, Fadi Al Machot, Kristian Hovde Liland, Habib Ullah

arXiv 2608.29923首次发表:更新:

发表机构

Norwegian University of Life Sciences (NMBU); Tulane University; University of Texas at Dallas(挪威生命科学大学; 杜兰大学; 德克萨斯大学达拉斯分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对开放词汇语义分割中视觉-语言模型在持续测试时分布偏移下的对齐脆弱问题,提出DAF框架,在多数据集上实现mIoU显著提升且鲁棒性良好。

AI 中文摘要

开放词汇语义分割(OVSS)依赖视觉-语言对齐来识别任意文本定义的类别,但这种对齐在持续测试时分布偏移下较为脆弱。我们的诊断分析显示:熵最小化会引发补丁级类别崩溃,持续更新会侵蚀视觉-语言对齐,低偏移样本产生的冗余梯度会浪费计算资源。我们提出了Diversify, Anchor, and Filter(DAF)稳定框架,该框架在基于熵的适应基础上,加入了抵抗崩溃的边际多样性损失、约束相对于冻结源模型的特征漂移的跨模态锚点一致性损失,以及跳过低价值反向传播以抵消部分源锚点开销的特征显著性过滤。我们在涵盖自然场景、自动驾驶、水下图像和遥感的5个数据集及其损坏变体上进行了评估。在评估的持续偏移场景中,DAF在熵最小化失效的情况下仍保持稳定,相比源模型,其在Pascal VOC20-C上的mIoU提升超过8个点,在LoveDA上提升超过9个点,在Foggy Cityscapes上提升超过3个点,且对激进的适应和学习率选择具有鲁棒性。

英文摘要

Open-vocabulary semantic segmentation (OVSS) relies on vision-language alignment to recognize arbitrary text-defined categories, yet this alignment is fragile under continual test-time distribution shift. Our diagnostic analysis reveals that entropy minimization drives patch-level class collapse, continual updates erode vision-language alignment, and redundant gradients from low-shift samples waste computation. We propose Diversify, Anchor, and Filter (DAF), a stabilization framework that augments entropy-based adaptation with a marginal diversity loss that resists collapse, a cross-modal anchor consistency loss that constrains feature drift relative to a frozen source model, and feature salience filtering that skips low-value backward passes to offset part of the source-anchor overhead. We evaluate on five datasets spanning natural scenes, autonomous driving, underwater imagery, and remote sensing with their corrupted variants. Across the evaluated continual shifts, DAF remains stable where entropy minimization collapses, improving mIoU by over 8 points on Pascal VOC20-C, over 9 points on LoveDA, and over 3 points on Foggy Cityscapes compared to the source model, and is robust to aggressive adaptation and learning rate choices.

Commentsunder review. code available at https://github.com/chandlerbing65nm/DAF.git

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑