面向计算高效的Transformer适配的电路微调
Circuit Fine-Tuning for Compute-Efficient Transformer Adaptation
浏览论文内容
中文总结 AI 辅助
提出电路微调(CFT)框架,通过电路发现选择ViT待微调子图,仅微调该子图,可大幅降低训练FLOPs与时间,在多类视觉任务上验证了有效性。
中文摘要 AI 辅助
参数高效微调(PEFT)已成为将视觉Transformer(ViTs)适配到下游任务的事实标准。尽管参数数量一直是PEFT中主要的效率指标,但它并不等同于计算效率:参数稀疏的方法每一步仍可能产生全模型训练成本,且通常需要较长的训练周期才能达到峰值精度。我们提出电路微调(Circuit Fine-Tuning, CFT),这是一种计算高效的框架,它利用电路发现(通常用于解释已训练模型的技术)在训练前选择要微调的模块。与传统归因是针对已训练的任务头不同,我们将其针对近零初始化的探测头进行归因,这能分离主干网络对目标分布的响应,而非特定分类器的偏好。随后,CFT仅对恢复的子图进行微调。CFT无需学习率预热,平均约20个epoch即可达到峰值精度,而强大的PEFT基线方法需要44至96个epoch,从而减少2.3至6.6倍的训练浮点运算量(FLOPs),并减少多达16倍的挂钟时间,同时不增加任何参数和推理操作。在标准视觉迁移基准(VTAB-1k)、分层主干网络(Swin)、领域偏移的医学影像(CBIS-DDSM)以及视觉语言模型(Gemma-3在CUB-200上)上的实验验证了CFT的有效性。代码可在该https URL获取。
英文摘要
Parameter-Efficient Fine-Tuning (PEFT) has become the de facto standard for adapting Vision Transformers (ViTs) to downstream tasks. While parameter count has been the dominant efficiency metric in PEFT, it does not imply \textit{compute efficiency}: parameter-sparse methods can still incur full-model training cost per step, and typically need long schedules to reach peak accuracy. We introduce Circuit Fine-Tuning (CFT), a compute-efficient framework that uses circuit discovery---conventionally used to explain trained models---to select modules for fine-tuning before training. Whereas attribution is conventionally formulated against a trained task head, we formulate it against a near-zero-initialized probe head, which isolates the response of the backbone to the target distribution rather than the preferences of a particular classifier. CFT then fine-tunes only the recovered subgraph. CFT needs no learning-rate warmup and reaches peak accuracy in ${\sim}20$ epochs on average---versus $44$--$96$ for strong PEFT baselines---yielding $2.3$--$6.6\times$ fewer training FLOPs and up to $16\times$ less wall-clock time, while adding zero parameters and no inference operations. Experiments across a standard visual transfer benchmark (VTAB-1k), hierarchical backbones (Swin), domain-shifted medical imaging (CBIS-DDSM), and a vision-language model (Gemma-3 on CUB-200) demonstrate the effectiveness of CFT. Code is available at https://github.com/UriKialy/CFT