部分之和大于整体:用于高效训练多策略大语言模型的自动任务排序
The Parts Are Greater Than the Sum: Automated Task Sequencing for Efficient Training of Multi-Policy LLMs
浏览论文内容
中文总结 AI 辅助
该研究针对多策略大语言模型训练中共享优化空间的干扰问题,提出自动多策略PEFT框架,通过任务分组排序组织独立QLoRA路径,在TRACE基准取得44.78的最佳性能。
中文摘要 AI 辅助
参数高效微调(PEFT)通常使用单个共享低秩适配器(LoRA)来适配大语言模型。当适配异构任务序列时,这个共享优化空间常受干扰,导致迁移效果差和灾难性遗忘。现有方法主要通过增加参数容量或组合多个适配器来提升适配器表达能力,但仍依赖共享优化路径。本文提出一种用于大语言模型参数高效微调的优化路径组织框架,实现为自动多策略PEFT架构。具体而言,在固定参数预算下,通过任务分组和任务排序自动组织优化兼容的适配路径,这些路径被实现为独立的量化低秩适配器(QLoRA),使异构任务能在解耦的适配空间中优化,同时保留兼容任务间的正向迁移。在TRACE基准上的实验表明,从传统单策略PEFT到多策略PEFT,性能持续提升,所提自动多策略框架在相同可训练容量下达到44.78的最佳性能,这说明优化路径组织比单纯增加适配器容量对异构参数高效微调更有效。
英文摘要
Parameter-Efficient Fine-Tuning (PEFT) commonly adapts large language models using a single shared Low-Rank Adapter (LoRA). This shared optimization space often suffers from interference when adapting heterogeneous task sequences, leading to poor transfer and catastrophic forgetting. Existing approaches mainly improve adapter expressiveness by increasing parameter capacity or composing multiple adapters, yet they still rely on a shared optimization path. In this paper, we propose an optimization-path organization framework for parameter-efficient fine-tuning of large language models, implemented as an automatic multi-policy PEFT architecture. Specifically, optimization-compatible adaptation paths are automatically organized through task grouping and task sequencing under a fixed parameter budget. The organized optimization paths are implemented as independent Quantized Low-Rank Adapters (QLoRA), enabling heterogeneous tasks to be optimized in decoupled adaptation spaces while preserving positive transfer among compatible tasks. Experiments on the TRACE benchmark demonstrate that performance consistently improves from conventional single-policy PEFT to multi-policy PEFT, with the proposed automatic multi-policy framework achieving the best performance of 44.78 under the same trainable capacity. This suggests that optimization-path organization is more effective than simply increasing adapter capacity for heterogeneous parameter-efficient fine-tuning.