发表机构
University of Amsterdam(阿姆斯特丹大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究将流形约束超连接作为参数高效微调新方法,包裹冻结主干。发现其与预训练设置作用不同,固定残差混合矩阵有益。独立时不比LoRA优,但mHC+LoRA组合在特定规模下有优势,确定残差路由是有前景的新型PEFT轴。
AI 中文摘要
大多数参数高效微调(PEFT)方法调整权重或激活,使关键的Transformer组件之一(残差连接)不变。本文研究了作为一种新型PEFT方法的流形约束超连接(mHC),它是残差连接的推广,用学习到的残差路由模块包裹冻结的OLMo-2主干。研究发现mHC能微调冻结的Transformer,但其作用与原始预训练设置有根本不同,固定残差混合矩阵为单位矩阵通常能提高性能。作为独立方法,mHC并不总是优于LoRA,但在匹配的可训练参数预算下,mHC+LoRA组合能改善语言建模损失,并在1B和7B规模上显示出任务相关的基准收益。总体而言,结果表明残差路由是一个独特且有前景的新型PEFT轴。
英文摘要
Finetuning methods for foundation models usually change weights, prompts, or hidden states, while leaving the residual topology fixed. We ask whether residual topology itself can become a finetuning object. To study this, we adapt manifold-constrained hyper-connections (mHC), recently introduced for pre-training, to frozen-backbone finetuning. mHC turns a Transformer into an input-dependent multi-stream residual architecture, routing representations through multiple streams at every sub-layer. Across mHC variants, we find that dynamic residual routing can finetune Transformers, but that its role differs from pre-training: by preserving stream mixing and learning only how sub-layers access streams, loss is improved and trainable parameters are reduced. Overall, our results identify residual routing as a promising architectural axis for efficient finetuning of foundation models.
CommentsNeurIPS 2026, AXIOM: Foundations of Efficient Deep Learning workshop