arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TopoTuner:大语言模型的拓扑微调

TopoTuner: Topological Finetuning of Large Language Models

Abdulkadir Erol, Yash Mahajan, Vepaul Hariprashad, Baha Rababah, Santu Karmaker, Cuneyt G. Akcora, Mubarak Shah

arXiv 2607.16637首次发表:更新:

发表机构

University of Central Florida; University of Manitoba(中佛罗里达大学; 曼尼托巴大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对大语言模型微调成本高及LoRA不足的问题,提出TopoTuner框架,通过拓扑引导选择性冻结注意力投影矩阵,从源数据集学习可复用冻结配置文件,在多模型上效果优于LoRA且大幅减少训练时间和参数更新量。

AI 中文摘要

完全微调仍然是适应预训练大语言模型的有效方法,但它会更新所有权重且成本高昂。LoRA减少了可训练参数的数量,但没有直接回答在适应过程中哪些预训练组件应被训练以及哪些可以冻结。我们引入了TopoTuner,这是一个用于选择性冻结注意力投影矩阵的拓扑引导微调框架。该方法将每个投影矩阵视为行云,并使用持久图之间的Wasserstein距离来衡量其在微调过程中的拓扑变化。TopoTuner从源数据集学习可重复使用的冻结配置文件,并将其转移以在域外数据集上高效微调模型,评估特定任务的拓扑漂移是否能跨问答和情感分析任务泛化。在多个模型上,TopoTuner与完全微调具有竞争力,同时仅训练1-2%的模型参数,在9个模型-数据集设置中的7个中优于LoRA,可改变高达39.57%的投影参数。它还减少了训练时间,相对于完全微调平均减少20.4%,相对于LoRA减少5.5%。TopoTuner为可重复使用的冻结配置文件开辟了新方向,即在一个数据集上学习的微调行为可跨多个任务共享。

英文摘要

Full fine-tuning remains a strong way to adapt pretrained LLMs, but it updates all weights and can be expensive. LoRA reduces the number of trainable parameters, but it does not directly answer which pretrained components should be trained and which can be frozen during adaptation. We introduce TopoTuner, a topology-guided fine-tuning framework for selective freezing of attention projection matrices. \method treats each projection matrix as a row cloud and uses Wasserstein distances between persistence diagrams to measure how its topology changes during fine-tuning. TopoTuner learns a reusable freezing profile from a source dataset and transfers it to efficiently fine-tune models on out-of-domain datasets, evaluating whether task-specific topological drift generalizes across question answering and sentiment analysis tasks. Across LLaMA-3.1-8B, Mistral-7B-v0.3, and Qwen3-8B-Base, TopoTuner is competitive with full fine-tuning while training only 1-2\% of the model parameters, and outperforms LoRA in 7 out of 9 model-dataset settings, which can change up to 39.57\% of the projection parameters. Along with minimized updates, TopoTuner reduces training time by 20.4\% relative to full fine-tuning and 5.5\% relative to LoRA on average. TopoTuner opens a new direction for reusable freezing profiles, where fine-tuning behavior learned on one dataset can be shared across multiple tasks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑