arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

直接自进化优化:无需挑战者训练即可进化大语言模型

Direct Self-Evolving Optimization: Evolving LLMs without Challenger Training

Yuyang Deng, Yu Wang, Jiayun Wang

arXiv 2609.34279首次发表:更新:

发表机构

Accenture, Center for Advanced AI; Georgia Institute of Technology(埃森哲,高级人工智能中心; 佐治亚理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出直接自进化优化(DEO),通过求解器引导的任务采样替代挑战者训练,实现大语言模型的高效自进化,在保持推理性能的同时显著降低训练时间。

AI 中文摘要

自进化语言模型通过生成任务并从自身反馈中学习来提升性能,但调整任务生成器通常需要单独的挑战者训练循环。我们能否在不显式训练挑战者的情况下,生成适应于当前求解器的任务?我们提出了直接自进化优化(DEO),该方法用求解器引导的任务采样取代了挑战者参数更新。KL正则化的挑战者目标定义了对固定基础任务分布的指数倾斜。DEO将此分布作为采样目标:一个冻结的大语言模型生成并变异任务,求解器对任务进行评分,近似Metropolis选择规则用于优化训练池。仅训练求解器。理论上,对于从倾斜分布精确采样的理想化变体,在正则性、局部梯度支配性和初始化条件下,我们证明DEO能学习到分布鲁棒的推理能力。实验中,DEO在推理性能上与R-Zero相当,同时减少了超过50%的墙钟训练时间,并相比无游走消融提高了推理准确性。用冻结的仅API大语言模型替换任务生成器进一步提升了局部求解器,展示了移除挑战者训练所实现的能力。

英文摘要

Self-evolving language models improve by generating tasks and learning from their own feedback, but adapting the task generator often requires a separate challenger-training loop. Can we generate tasks adapted to the current solver without explicitly training a challenger? We introduce \textbf{D}irect Self-\textbf{E}volving \textbf{O}ptimization (DEO), which replaces challenger parameter updates with solver-guided task sampling. The KL-regularized challenger objective defines an exponential tilt of a fixed base task distribution. DEO uses this distribution as a sampling target: a frozen LLM generates and mutates tasks, the solver scores them, and an approximate Metropolis selection rule refines the training pool. Only the solver is trained. Theoretically, for an idealized variant that samples exactly from the tilted distribution, and under regularity, local gradient-dominance, and initialization conditions, we show that DEO learns distributionally robust reasoning ability. In experiments, DEO achieves reasoning performance competitive with R-Zero while using over $50\%$ less wall-clock training time, and improves reasoning accuracy over a no-walk ablation. Replacing the task generator with a frozen API-only LLM further improves the local solver, illustrating a capability enabled by removing challenger training.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑