发表机构
Aalborg University; Pioneer Centre for Artificial Intelligence(奥尔堡大学; 人工智能先锋中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出BiSSLight,结合M-FAC隐式梯度近似与LoRA参数高效微调,高效解决自监督与下游任务目标错位问题,在更大规模模型上显著提升下游性能并大幅降低计算成本。
AI 中文摘要
自监督预训练学习到的表示可广泛迁移至多种下游任务,然而直接微调可能因自监督目标与下游任务目标之间的错位而表现次优,甚至可能损害对下游任务有益的预训练特征。BiSSL框架通过引入一个过渡训练阶段解决了这一问题,该阶段被构建为双层优化问题,其中下游任务目标引导自监督学习过程,以精炼预训练表示,从而更好地促进后续微调。然而,BiSSL依赖传统的双层优化求解技术,其昂贵的隐式超梯度近似使得该方法对当代模型架构越来越不实用。为使其高效且可扩展,我们提出了BiSSLight,它将基于M-FAC的隐式梯度近似与通过LoRA实现的参数高效微调相结合,从而能够在先前不切实际的更大规模上高效应用。在多个下游任务和当代模型架构上的评估表明,BiSSLight持续提升下游性能,且随着模型规模增大,尽管基线更强,性能提升反而更为显著。该方法计算效率极高,在ViT-H骨干网络上相比其前身将计算时间减少了超过十倍。
英文摘要
Self-supervised pretraining learns representations that are broadly transferable across downstream tasks, yet direct fine-tuning can be suboptimal due to misalignment between self-supervised and downstream task objectives, potentially degrading pretrained features beneficial to the downstream task. The BiSSL framework addressed this by introducing a transitional training stage formulated as a bilevel optimization problem, in which the downstream task objective guides the self-supervised learning process in refining pretrained representations to better facilitate subsequent fine-tuning. However, BiSSL relies on conventional bilevel optimization solving techniques whose costly implicit hypergradient approximations render the method increasingly impractical for contemporary model architectures. To make it efficient and scalable, we introduce BiSSLight, which combines M-FAC-based implicit gradient approximation with parameter-efficient fine-tuning via LoRA, enabling efficient application at larger scales that were previously impractical. Evaluation across multiple downstream tasks and contemporary model architectures shows that BiSSLight consistently improves downstream performance, with gains becoming more pronounced as model size increases despite stronger baselines. The method is highly computationally efficient, reducing computation time by more than a factor of ten compared to its predecessor on a ViT-H backbone.