arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2310.12100cs.CLcs.AIcs.CVcs.LGcs.MM

非侵入式自适应:以输入为中心的参数高效微调用于多功能多模态建模

Non-Intrusive Adaptation: Input-Centric Parameter-efficient Fine-Tuning for Versatile Multimodal Modeling

  • Google(谷歌)

机构由 AI 辅助整理,请以论文原文为准。

Yaqing Wang, Jialin Wu, Tanmaya Dabral, Jiageng Zhang, Geoff Brown, Chun-Ta Lu, Frederick Liu, Yi Liang, Bo Pang, Michael Bendersky, Radu Soricut

更新

AI总结:

本文提出非侵入式PEFT技术AdaLink,仅调整模型外部输入参数,在纯文本和多模态任务上取得了与侵入式PEFT(LoRA)及全模型微调相媲美的性能。

AI中文摘要:

大型语言模型(LLMs)和视觉语言模型(VLMs)通过将参数量从O(10^9)扩展到O(10^{12})甚至更高,在广泛的任务上展现出卓越性能。这些庞大的规模使得在给定感兴趣任务时,无法适应并部署完全专用的模型。参数高效微调(PEFT)作为解决此类大型模型自适应与部署挑战的有前景方向应运而生。我们将PEFT技术分为两类:侵入式和非侵入式。侵入式PEFT技术直接改变模型的内部架构。尽管更灵活,但它们为训练和部署引入了显著复杂性。非侵入式PEFT技术保持内部架构不变,仅调整模型外部参数,例如输入的嵌入。在本工作中,我们将AdaLink描述为一种非侵入式PEFT技术,在多种任务上与SoTA侵入式PEFT(LoRA)和全模型微调(FT)相比取得了有竞争力的性能。我们使用纯文本和多模态任务进行评估,实验考虑了参数量扩展和训练机制(有无指令微调)。

英文摘要:

Large language models (LLMs) and vision language models (VLMs) demonstrate excellent performance on a wide range of tasks by scaling up parameter counts from O(10^9) to O(10^{12}) levels and further beyond. These large scales make it impossible to adapt and deploy fully specialized models given a task of interest. Parameter-efficient fine-tuning (PEFT) emerges as a promising direction to tackle the adaptation and serving challenges for such large models. We categorize PEFT techniques into two types: intrusive and non-intrusive. Intrusive PEFT techniques directly change a model's internal architecture. Though more flexible, they introduce significant complexities for training and serving. Non-intrusive PEFT techniques leave the internal architecture unchanged and only adapt model-external parameters, such as embeddings for input. In this work, we describe AdaLink as a non-intrusive PEFT technique that achieves competitive performance compared to SoTA intrusive PEFT (LoRA) and full model fine-tuning (FT) on various tasks. We evaluate using both text-only and multimodal tasks, with experiments that account for both parameter-count scaling and training regime (with and without instruction tuning).

↑