发表机构
Korea University; Korea Institute of Industrial Technology; Seoul National University(韩国大学; 韩国产业技术研究院; 首尔国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对视觉基础模型自适应的挑战,提出低秩卷积自适应框架 LoCA,通过解耦通道与空间自适应,引入低秩通道自适应并优化空间基,在多视觉任务中保留预训练空间先验且性能优异。
AI 中文摘要
预训练的视觉基础模型(VFM)为各种下游任务提供强大的视觉表示。VFM 自适应的关键挑战源于完全微调的高昂成本和灾难性遗忘。为解决此问题,低秩自适应(LoRA)已成为参数高效微调(PEFT)的主流范式。然而,LoRA 通常是为 2D 矩阵参数化的变压器自注意力层设计的。由于卷积核在 4D 张量中固有地耦合空间和通道信息,将它们强制转换为单一的 2D 矩阵会破坏固有的空间拓扑结构。本文提出了低秩卷积自适应(LoCA),这是一种卷积感知的 PEFT 框架,通过解耦通道和空间自适应来解决空间通道纠缠问题。LoCA 引入了用于密集跨通道混合的低秩通道自适应,并通过奇异值分解(SVD)优化从预训练核中提取的空间基。实验结果表明,LoCA 保留了预训练的空间先验,并在细粒度分类、领域通用语义分割和生成基准测试中取得了有竞争力或领先的性能。
英文摘要
Pre-trained Vision Foundation Models (VFMs) provide strong visual representations for diverse downstream tasks. The key challenge of VFM adaptation stems from the prohibitive costs of full fine-tuning and catastrophic forgetting. To address this, Low-Rank Adaptation (LoRA) has emerged as the prevailing paradigm for Parameter-Efficient Fine-Tuning (PEFT). However, LoRA is typically designed for transformer self-attention layers parameterized by 2D matrices. Since convolutional kernels inherently couple spatial and channel information within a 4D tensor, forcing them into a monolithic 2D matrix disrupts the inherent spatial topology. In this paper, we propose Low-Rank Convolutional Adaptation (LoCA), a convolution-aware PEFT framework that addresses spatial-channel entanglement by decoupling channel and spatial adaptation. LoCA introduces a low-rank channel adaptation for dense cross-channel mixing and refines spatial bases extracted from pre-trained kernels via Singular Value Decomposition (SVD). Experimental results show that LoCA preserves pre-trained spatial priors and achieves competitive or state-of-the-art performance across fine-grained classification, domain-generalized semantic segmentation, and generative benchmarks.
CommentsAccepted by ECCV 2026