发表机构
Northeastern University(东北大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出LS-B方法,利用fMRI脑读出信号指导持续学习者的适配器放置,无需搜索即可在中间层实现高效专业化,性能接近全适配器且存储大幅降低。
AI 中文摘要
持续学习者在预训练视觉Transformer的每个块中保留任务特定适配器时,其存储量会随任务数量线性增长;仅在少数块中保留任务特定适配器可抑制这种增长,但引发了这些适配器应放置于何处的问题。我们从两个角度研究这一问题。在算法层面,训练所有连续的四个块放置方案产生倒U形结果:最终准确率在中间深度达到峰值,且变化幅度高达3.5个百分点(pp),而基于权重谱或激活统计的低成本标准则倾向于最深的块。从神经科学角度,视觉皮层的中级阶段可塑性和层级组织促使我们探究:能否通过在学习器外部进行的测量来指导层专业化,而无需进行放置搜索。LS-B通过冻结的十二个人类视觉区域fMRI编码模型观察前几个任务,并将任务特定容量一次性分配给那些读出值在不同任务间变化最大(相对于其稳定结构)的块。在三种ViT-B/16骨干网络上,LS-B产生稳定且骨干特定的分配。在具有放置搜索的两种骨干网络(AugReg和iBOT)上,所选块与搜索确定的中间深度区域重叠。在匹配的存储和观察预算下,所选块优于最浅和最深的四个块配置。在Split ImageNet-R上,LS-B使用全BiLoRA适配器存储的60%,同时保持在其最终准确率的1.5个百分点以内。该分配无需标签或反向传播,运行时增加低于0.6%,并表现出骨干特定的皮层特征。
英文摘要
Continual learners that keep a task-specific adapter in every block of a pre-trained vision transformer accumulate storage linearly with the number of tasks; keeping task-specific adapters in only a few blocks curbs this growth but raises the question of where to place them. We investigate this question from two perspectives. Algorithmically, training all contiguous four-block placements yields an inverted U: final accuracy peaks at intermediate depth and varies by up to 3.5 percentage points (pp), while inexpensive criteria based on weight spectra or activation statistics favor the deepest blocks. From neuroscience, the hierarchical organization and intermediate-stage plasticity of the visual cortex motivate us to ask whether a measurement taken outside the learner can guide layer specialization without placement search. LS-B observes the first tasks through a frozen fMRI encoding model of twelve human visual areas and commits task-specific capacity once to the blocks whose readouts vary most across tasks relative to their stable structure. Across three ViT-B/16 backbones, LS-B yields stable, backbone-specific allocations. On the two backbones with placement search, AugReg and iBOT, the selected blocks overlap the intermediate-depth region identified by search. Under matched storage and observation budgets, the selected blocks outperform the shallowest and deepest four-block configurations. On Split ImageNet-R, LS-B uses 60% of full-BiLoRA adapter storage while remaining within 1.5 pp of its final accuracy. The allocation requires no labels or backpropagation, adds under 0.6% runtime, and exhibits backbone-specific cortical signatures.
Comments21 pages, 12 figures