DIVE:通过多样性驱动的技能进化解锁冻结语言模型的自我提升
DIVE: Unlocking Self-Improvement in Frozen Language Models Through Diversity-Driven Skill Evolution
浏览论文内容
中文总结 AI 辅助
DIVE是一种多样性驱动的无参数框架,可让冻结LLM从任务经验和验证器反馈中进化出自然语言技能,在多项推理任务上优于现有方法,还能实现模型间技能迁移,让小模型性能匹配或超越大模型。
中文摘要 AI 辅助
大型语言模型(LLM)若不进行参数更新,便无法保留部署后的经验。我们提出DIVE,这是一种多样性驱动的框架,能让冻结的LLM通过从任务经验和验证器反馈中进化出持久的自然语言技能来实现自我提升。这些技能编码了可复用的推理流程、验证策略、常见失败模式以及输出约束,且由同一基础模型执行和修订,无需访问教师模型。由于自然语言技能进化是一个随机非凸搜索过程,优化单一技能轨迹可能会过拟合采样经验或收敛到次优解。DIVE通过从自举经验中独立进化多个技能种群、通过多样化变换自适应优化这些种群、并联合选择互补技能集来缓解这种优化方差。在六个数学和逻辑推理任务以及多个模型家族上,DIVE始终优于现有推理方法、提示优化方法、技能开发框架和基于记忆的基线。它能从累积经验中实现快速自我提升,与SFT、GRPO等基于参数的方法以及使用GEPA的提示优化相比,仅需更少的rollout就能获得显著更大的性能提升。此外,生成的技能可跨模型规模和家族迁移,使GPT-5-nano等较小模型在常规提示下能匹配或超越更大的对应模型(即GPT-5)。这些结果表明,多样性驱动的技能进化是一种有效、可解释且无参数的LLM自我提升方法。
英文摘要
Large language models (LLMs) cannot retain post-deployment experience without parameter updates. We introduce DIVE, a diversity-driven framework that enables frozen LLMs to improve by evolving persistent natural-language skills from task experience and verifier feedback. These skills encode reusable reasoning procedures, verification strategies, common failure modes, and output constraints and are both executed and revised by the same underlying model without access to a teacher model. Since natural-language skill evolution is a stochastic, non-convex search process, optimizing a single skill trajectory can overfit to sampled experience or converge to a suboptimal solution. DIVE mitigates this optimization variance by independently evolving multiple skill populations from bootstrapped experience, adaptively refining them through diverse transformations, and jointly selecting a complementary set of skills. Across six mathematical and logical reasoning tasks and multiple model families, DIVE consistently outperforms existing reasoning methods, prompt-optimization approaches, skill-development frameworks, and memory-based baselines. It achieves rapid self-improvement from accumulated experience, obtaining substantially larger performance gains with fewer rollouts than parameter-based methods such as SFT and GRPO, and prompt optimization with GEPA. Further, the resulting skills transfer across model scales and families, enabling smaller models such as GPT-5-nano to match or outperform larger counterparts, i.e., GPT-5, under conventional prompting. These results establish diversity-driven skill evolution as an effective, interpretable, and parameter-free approach to LLM self-improvement.