位置感知层查询用于视觉语言模型中的测试时训练
Position Aware Layer Queries for Test Time Training in Vision Language Models
- Institute of Artificial Intelligence(人工智能研究所)
- University of Central Florida(中佛罗里达大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
提出位置感知层查询网络(LQN),通过单次前向传播和位置感知蒸馏及位置一致性正则化,高效适应冻结的视觉语言模型,提升分布外和细粒度分类性能。
中文摘要 AI 辅助
测试时训练(TTT)在传统微调不可行时,使模型适应到来的测试样本(例如分布外(OOD)样本)。现有的针对视觉语言模型(VLM)的TTT方法从多个增强视图创建监督信号,每个视图都需要通过整个VLM进行前向(通常还有反向)传播,从而产生大量计算成本。我们观察到,一次前向传播中所有中间层输出已经比所有增强的最终嵌入提供了多得多的信号。我们引入了层查询网络(LQN),这是一种轻量级方法,可以通过一个小型模型(学生)在VLM的单次前向传播中适应冻结的VLM(教师)。LQN使用位置感知蒸馏(PAD)通过查询中间标记的空间坐标来模仿教师VLM的中间层空间标记。LQN还依赖位置一致性正则化(LCR),这是一种自监督技术,用O(1)坐标采样替代昂贵的O(H x W)图像增强。整合这些,LQN i)在OOD ImageNet上将零样本CLIP ViT-B/16的Top-1准确率提高了9.8%,ii)在细粒度分类上比之前最好的GS-Bias高出3.9%的Top-1,iii)在CLIP ResNet-50上比TPS收敛更快(47分钟对比55分钟),iv)将适应推广到SigLIP、EVA-CLIP和CoCa等VLM,以及MLP、ResNet、VGG等轻量级学生网络,v)扩展到全景、实例和语义分割。
英文摘要
Test-Time Training (TTT) adapts models to incoming test samples (e.g. out-of-distribution, (OOD)) when conventional fine-tuning is infeasible. Existing TTT methods for Vision-Language Models (VLMs) create supervision from several augmented views, each requiring forward (and often backward) passes through the entire VLM, incurring substantial computational cost. We observe that one forward pass with all the intermediate layer outputs already yields far more signal than the final embedding from all augmentations. We introduce Layer Query Network (LQN), a lightweight approach that can adapt a frozen VLM (teacher) in a single forward pass of the VLM via a small model (student). LQN uses Position-Aware Distillation (PAD) to mimic the teacher VLM's intermediate-layer spatial tokens by querying spatial coordinates of intermediate tokens. LQN additionally relies on Location Consistency Regularization (LCR), a self-supervision technique, replacing expensive O(H x W) image augmentation with O(1) coordinate sampling. Integrating these, LQN i) adapts and improves zero-shot CLIP ViT-B/16 by 9.8% Top-1 on OOD ImageNet, ii) outperforms the previous best GS-Bias on fine-grained classification by 3.9% Top-1, iii) achieves faster convergence than TPS for CLIP ResNet-50 (47 mins vs 55 mins), iv) generalizes adaptation to VLMs like SigLIP, EVA-CLIP, and CoCa, and lightweight students like MLP, ResNet, VGG, and v) extends to panoptic, instance, and semantic segmentation.