arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于参数高效通用视觉适配的自路由张量适配器

Self-Routed Tensor Adapters for Parameter-Efficient Universal Visual Adaptation

Suraj Yadav

arXiv 2608.16384首次发表:更新:

发表机构

Indraprastha Institute of Information Technology Delhi(因德拉普拉斯信息技术学院德里分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出自路由张量适配器(SRTA),无需外部门控网络即可实现样本特定适配,在多领域视觉分类基准上,以更少参数达到与MoE式PEFT基线相当或更优的平均准确率。

AI 中文摘要

通用视觉表示需要适配机制,使其能跨异构领域适配,且不会将知识拆分为特定领域的模块。参数高效微调可高效适配冻结的视觉基础模型,但标准低秩适配器对所有输入使用固定子空间,当领域在风格、背景和语义上下文存在差异时,这种方式会受限。基于MoE的适配器通过多个专家路径提升专业化程度,但通常依赖外部路由器和大型专家库,会增加参数并使路由与适配分离。我们提出**自路由张量适配器(Self-Routed Tensor Adapters,SRTA)**,这是一种用于多领域视觉适配的紧凑框架。SRTA将每个输入投影到低秩空间,利用可学习的领域矩阵从该表示计算路由权重,并使用这些权重混合共享Tucker核的切片。这无需外部门控网络即可生成样本特定的适配矩阵,使共享视觉因子可重复使用,同时支持领域感知的专业化。为强化路径学习,我们引入渐进式深度加权路由目标,对适配器层的路由决策进行监督。在五个异构多领域视觉分类基准上,SRTA的平均准确率与MoE式参数高效微调(PEFT)基线相当或略高,同时使用的可训练参数显著更少。在秩为64时,SRTA在4领域设置中使用277万参数,而MoLoRA为952万;在6领域设置中,SRTA使用300万参数,MoLoRA为1431万。总体而言,SRTA为将视觉基础模型适配为通用多领域表示提供了有效的准确率-参数权衡方案。GitHub

英文摘要

Universal visual representations require adaptation mechanisms that adapt across heterogeneous domains without fragmenting knowledge into domain-specific modules. Parameter-efficient fine-tuning adapts frozen visual foundation models efficiently, but standard low-rank adapters use a fixed subspace for all inputs, which can be restrictive when domains differ in style, background, and semantic context. MoE-based adapters improve specialization through multiple expert pathways, but often rely on external routers and large expert banks, adding parameters and separating routing from adaptation. We propose \textbf{Self-Routed Tensor Adapters}, a compact framework for multi-domain visual adaptation. SRTA projects each input into a low-rank space, computes routing weights from this representation using a learnable domain matrix, and uses these weights to blend slices of a shared Tucker core. This produces a sample-specific adaptation matrix without an external gating network, allowing shared visual factors to be reused while supporting domain-aware specialization. To strengthen pathway learning, we introduce a progressive depth-weighted routing objective that supervises routing decisions across adapter layers. Across five heterogeneous multi-domain visual classification benchmarks, SRTA achieves competitive or slightly stronger average accuracy than MoE-style PEFT baselines while using substantially fewer trainable parameters. At rank 64, SRTA uses 2.77M parameters in the 4-domain setting compared with 9.52M for MoLoRA, and 3.00M in the 6-domain setting compared with 14.31M. Overall, SRTA offers an effective accuracy-parameter trade-off for adapting visual foundation models toward universal multi-domain representations. \href{https://github.com/surajyadav-research/SRTA}{GitHub}

CommentsAccepted at ECCV Workshop 2026 (Archival Track)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑