arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12679cs.AIcs.NE

超越最佳猜测:用进化策略提升大语言模型的解决方案覆盖率

Beyond the Best Guess: Improving LLM Solution Coverage with Evolution Strategies

  • The University of Texas at Austin(德克萨斯大学奥斯汀分校)
  • Cognizant AI Lab(高知特人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

Conor F. Hayes, Elliot Meyerson, Kajetan Schweighofer, Roberto Dailey, Babak Hodjat, Risto Miikkulainen, Xin Qiu

AI总结:

该研究针对LLM后训练中RL导致pass@k受限、解决方案覆盖率不足的问题,采用进化策略(ES)方法,提升了pass@k与解决方案覆盖率,在数学基准上取得更好结果,为相关领域后训练提供了更好基础。

AI中文摘要:

大语言模型(LLMs)正越来越多地应用于数学、科学等发现领域。通常的做法是将问题呈现给模型,并将其答案作为拟议的解决方案。然而,除了这个最佳猜测之外,增加测试时计算可以增强发现过程。在一个名为pass@k的过程中,模型被允许探索解决方案空间并生成多样化的候选解决方案。遗憾的是,通过强化学习(RL)对LLMs进行后训练的标准方法可能会限制pass@k:模型的输出分布会在高奖励输出周围收窄,导致解决方案覆盖率崩溃。替代方案是使用进化策略(ES),这是一种基于种群、无梯度的后训练方法,通过随机扰动直接在权重空间中优化。正如本文所示,ES实现了比RL始终更高的pass@k,并产生了更广泛的输出分布,具有更大的解决方案覆盖率。这种覆盖率反过来使得在标准数学基准等方面取得更好的结果成为可能。因此,ES为发现问题以及其他对多样化解决方案覆盖率至关重要的领域的后训练提供了更好的基础。

英文摘要:

Large Language Models (LLMs) are increasingly deployed in discovery domains such as math and science. The usual approach is to present the problem to the model and use its answer as the proposed solution. However, beyond this best guess, discovery can be enhanced by increasing test-time compute. In a process called pass@k, the model is allowed to explore the solution space and generate diverse candidate solutions. Unfortunately, the standard approach to post-training LLMs through Reinforcement Learning (RL) may limit pass@k: the model's output distribution narrows around high-reward outputs, causing the solution coverage to collapse. The alternative is to use Evolution Strategies (ES), a population-based, gradient-free post-training method that optimizes directly in weight space through random perturbations. As this paper shows, ES achieves consistently higher pass@k than RL and produces a broader output distribution with greater solution coverage. This coverage in turn makes it possible to achieve better results in e.g. standard math benchmarks. Thus, ES provides a better foundation for post-training in discovery problems and other domains where diverse solution coverage is critical.

↑