发表机构
Xidian University(西安电子科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对视觉基础模型适配指代图像分割时全量微调开销大、单向交互不足的问题,提出双向互惠学习框架,通过互惠注意力与门控适配器实现双向跨模态交互,以不到0.5%的参数更新取得最先进性能。
AI 中文摘要
视觉基础模型(VFMs)的最新进展已在各种单模态视觉任务中展现出卓越能力。然而,将视觉基础模型适配到指代图像分割(RIS)通常需要通过全量微调实现精确的视觉-语言对齐,这会产生大量的计算开销,并存在灾难性遗忘的风险。现有的参数高效微调(PEFT)方法虽然能以最小的训练成本实现安全的知识迁移,但它们主要在单个模态内独立运作,或仅专注于从语言到视觉的单向引导,忽视了渐进式的跨模态交互以及用于文本细化的视觉反馈。为解决这些局限性,我们提出了双向互惠学习(BRL),一种新颖的基于适配器的参数高效微调框架,该框架在冻结基础模型的token混合层和通道混合层内促进分层、双向的信息流动。具体而言,BRL引入了两个互补的轻量级模块。互惠注意力适配器(RAA)在token级别执行跨模态的查询-键交换,使视觉和语言token能够相互关注,以实现细粒度的空间定位。互惠门控适配器(RGA)在通道级别生成跨模态门控信号,允许来自一种模态的全局语义上下文自适应地重新校准另一种模态的通道激活。在RefCOCO、RefCOCO+和RefCOCOg基准上的大量实验证明了BRL相对于先前RIS方法的优越性,在实现最先进性能的同时,仅需更新不到0.5%的骨干网络参数。代码和模型将在此https URL发布。
英文摘要
Recent advances in vision foundation models (VFMs) have shown remarkable capabilities across diverse unimodal visual tasks. However, adapting VFMs to referring image segmentation (RIS) typically necessitates precise vision-language alignment via full fine-tuning, incurring substantial computational overhead and risking catastrophic forgetting. While existing parameter-efficient fine-tuning (PEFT) methods enable safe knowledge transfer with minimal training costs, they predominantly operate independently within individual modalities or focus exclusively on unidirectional guidance from language to vision, overlooking progressive cross-modal interaction and visual feedback for textual refinement. To address these limitations, we propose $\textbf{B}$idirectional $\textbf{R}$eciprocal $\textbf{L}$earning ($\textbf{BRL}$), a novel adapter-based PEFT framework that facilitates hierarchical, bidirectional information flow within both token-mixing and channel-mixing layers of frozen foundation models. Specifically, BRL introduces two complementary lightweight modules. The Reciprocal Attention Adapter (RAA) performs cross-modal query-key exchanges at the token level, enabling visual and linguistic tokens to mutually attend to each other for fine-grained spatial grounding. The Reciprocal Gate Adapter (RGA) generates cross-modal gating signals at the channel level, allowing global semantic context from one modality to adaptively recalibrate channel activations of the other. Extensive experiments on RefCOCO, RefCOCO+, and RefCOCOg benchmarks demonstrate the superiority of BRL over prior RIS methods, achieving state-of-the-art performance while requiring less than 0.5% backbone parameter updates. Code and models will be released at https://github.com/xiaoqiang-lu/BRL.
Comments16 pages, 8 figures