StraTune:自进化大语言模型技能中修订算子的自适应选择
StraTune: Adaptive Selection of Revision Operators for Self-Evolving LLM Skills
AI总结:
针对自进化大语言模型技能修订中固定算子效果不佳的问题,提出StraTune方法,通过冻结优化器按优化状态自适应选择修订算子,在多个基准上超越基线。
AI中文摘要:
大语言模型(LLMs)可以从执行反馈中学习可复用的文本技能,而无需更新其参数,但如何有效地决定如何修订这些技能仍是一个关键挑战。现有方法通常依赖于固定的修订算子、搜索策略以及在该策略下应用的修订形式。然而,我们观察到,没有任何单一的修订算子在所有任务中始终表现最佳,并且反复应用不合适的算子可能会限制进一步的改进。我们提出了StraTune(策略引导的技能调优),它让一个冻结的优化器大语言模型根据优化状态(定义为当前执行反馈以及先前策略和形式的记录结果)在每一轮选择修订算子。来自每个修订算子的候选技能会经过一次候选评估,该评估在小样本集上筛选增益和回归,并在更大样本集上进行验证,且每个结果都会被写回优化状态以供后续选择。在四个基准和两种大语言模型设置中,StraTune在大多数设置中优于所有五个基线。消融实验将增益归因于修订算子的自适应选择,因为固定、随机、预定和赌博机策略选择的得分均较低,并且使用小型目标大语言模型学习的技能也能提升更强的模型。代码和学习到的技能可在以下网址获取:此https URL。
英文摘要:
Large language models (LLMs) can learn reusable textual skills from execution feedback without updating their parameters, but effectively deciding how to revise these skills remains a key challenge. Existing methods typically rely on a fixed revision operator, a search strategy and the revision forms applied under it. However, we observe that no single revision operator consistently performs best across tasks, and repeatedly applying an unsuitable operator can limit further improvement. We propose StraTune (strategy-guided skill tuning), which lets a frozen optimizer LLM choose the revision operator at every round from the optimization state, which is defined as the current execution feedback together with the recorded outcomes of earlier strategies and forms. Candidate skills from every revision operator pass one candidate evaluation, which screens for gains and regressions on a small sample set and validates them on a larger one, and every outcome is written back to the optimization state for later choices. Across four benchmarks and two LLM settings, StraTune outperforms all five baselines in most settings. Ablations attribute the gains to the adaptive choice of the revision operator, since fixed, random, scheduled, and bandit strategy choices all score lower, and skills learned with a small target LLM also improve a stronger one. Code and learned skills are available at https://github.com/seai-lab/StraTune.