发表机构
University of Science, Viet Nam National University, Ho Chi Minh City; University of Dayton(胡志明市越南国立大学理学院; 代顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
LAS-CLIP提出一种轻量级适配器引导方法,冻结CLIP全部参数,通过MaskAdapter注入注意力偏置,实现区域级任务上的零样本能力保持,且参数极少、性能优于Alpha-CLIP。
AI 中文摘要
CLIP的视觉编码器仅生成全局图像表示,这限制了其在区域级任务中的使用。现有的适配方法依赖于视觉提示、输入掩码或编码器微调,每种方法都会损害预训练表示。我们提出LAS-CLIP,一种轻量级适配器引导方法,该方法保持CLIP的所有参数冻结。一个紧凑的MaskAdapter根据输入掩码生成每头、每层的注意力偏置,并将其注入到冻结的自注意力层中,将注意力引导至目标区域。关键的是,由于主干网络保持严格不变,当未提供掩码时,LAS-CLIP可无缝恢复为原始CLIP,保留其基础零样本能力。凭借约116K至145K的可训练参数以及在两块T4 GPU上的100K训练样本,LAS-CLIP在ImageNet-S零样本分类和RefCOCO指代表达理解任务上取得了与Alpha-CLIP相当或更优的结果,尽管后者在数百万样本上微调了其整个编码器。定性分析进一步证实了在错误掩码和下游生成场景下更强的表示保真度。我们的项目页面链接为https URL。
英文摘要
CLIP's visual encoder produces only global image representations, limiting its use in region-level tasks. Existing adaptations rely on visual prompting, input masking, or encoder fine-tuning, each compromising pre-trained representations. We propose LAS-CLIP, a Lightweight Adapter Steering approach that keeps every CLIP parameter frozen. A compact MaskAdapter generates per-head, per-layer attention biases from an input mask and injects them into the frozen self-attention layers, steering attention toward the target region. Crucially, because the backbone remains strictly untouched, LAS-CLIP seamlessly reverts to vanilla CLIP when no mask is provided, preserving its foundational zero-shot capabilities. With approximately 116K to 145K trainable parameters and 100K training samples on two T4 GPUs, LAS-CLIP achieves competitive or superior results compared to Alpha-CLIP on ImageNet-S zero-shot classification and RefCOCO referring expression comprehension, despite the latter fine-tuning its entire encoder on millions of samples. Qualitative analysis further confirms stronger representational fidelity under incorrect masks and in downstream generation.