arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

策略积累与引导执行用于自动化大语言模型微调

Strategy Accumulation and Guided Execution for Automated LLM Fine-Tuning

Haoran Zhao, Wei Du, Dingwen Yang, Jixuan Huang, Junlin Shang, Lingyong Fang, Ya Guo, Tao Gui, Qi Zhang, Xuanjing Huang

arXiv 2609.22257首次发表:更新:

发表机构

Fudan University; Ant Group; Shanghai Jiaotong University(复旦大学; 蚂蚁集团; 上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对自动化微调系统无状态导致重复搜索的问题,提出SAGE两阶段框架,通过多智能体MCTS探索和蒸馏积累经验,在九个未见任务上将平均相对改进从3.2%提升至15.6%。

AI 中文摘要

生成特定任务的大语言模型需要通过实验发现有效的训练策略。自动化微调系统已使这种实验变得可行,且所需人工努力大大减少。然而,这些系统是无状态的:每次搜索一旦结束,其发现的策略、数据集洞察和超参数结果便被丢弃。每个新任务都必须从冷启动重复这种代价高昂的搜索。为解决此问题,我们提出策略积累与引导执行(SAGE),一个两阶段框架,使自动化微调搜索具有累积性。在第一阶段,一个多智能体管道执行基于蒙特卡洛树搜索的探索。一个并行蒸馏智能体提取任务特定的探索记录和置信度评分的跨任务洞察,这些共同构成一个结构化的经验库。在第二阶段,SAGE从该库中检索相关经验,并选择适用的部分来指导新任务上的训练。我们在九个未见任务上评估SAGE,涵盖单类别和跨类别设置。在单轮执行中,SAGE积累的经验将相对于基线的平均相对改进从3.2%提升至15.6%,相比没有该经验的同一管道,提升了12.4个百分点。这些结果表明,持久化的策略经验为未见任务上的自动化微调提供了有效指导。

英文摘要

Producing task-specific large language models requires discovering effective training strategies through experimentation. Automated fine-tuning systems have made this experimentation feasible with far less manual effort. However, these systems are stateless: each search discards its discovered strategies, dataset insights, and hyperparameter findings once it ends. Every new task must then repeat this costly search from a cold start. To address this, we propose Strategy Accumulation and Guided Execution (SAGE), a two-stage framework that makes automated fine-tuning search cumulative. In the first stage, a multi-agent pipeline performs Monte Carlo Tree Search-based exploration. A parallel Distillation Agent extracts task-specific exploration records and confidence-scored cross-task insights, which together constitute a structured experience repository. In the second stage, SAGE retrieves relevant experience from this repository and selects what applies to guide training on the new task. We evaluate SAGE on nine unseen tasks spanning both single- and cross-category settings. In single-round execution, SAGE's accumulated experience raises the average relative improvement over baseline from 3.2% to 15.6%, a 12.4-percentage-point gain over the same pipeline without it. These results show that persistent strategy experience provides effective guidance for automated fine-tuning on unseen tasks.

Comments24 pages, 4 figures, 10 tables, including supplementary material

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑