arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过异构编辑重组克服大语言模型驱动的程序优化中的最弱链效应

Overcoming the Weakest-Link Effect in LLM-Driven Program Optimization via Heterogeneous Edit Recombination

Jingwen Fu, Zhen Liu, Yuhan Liu, He Zhang, Nanning Zheng

arXiv 2607.28947首次发表:更新:

AI 中文总结

本研究针对LLM程序优化的最弱链效应,提出HERO异构编辑重组优化器,可生成多样非重叠原子编辑并结合评估器分数组合改进,在多领域实现更高分数、更快收敛与更少token消耗。

AI 中文摘要

大语言模型(LLM)正越来越多地通过在程序空间中搜索来解决复杂问题,为可自然表示和求解为程序的科学问题提供了通用范式。尽管近期取得了进展,但为候选程序确定有效的优化方向仍然具有挑战性。与自动微分类似,现有方法通常使用文本“梯度”引导搜索:即表示为文本编辑的一阶更新方向,这些梯度要么从先前评估的程序中推断,要么从LLM生成的隐式程序-分数映射反馈中推断。然而,随着程序-分数映射变得更复杂,这些估计的可靠性会不断降低,限制了它们的实际应用。我们认为显式梯度对于有效的程序优化并非必不可少,LLM可利用其先验知识直接从当前程序中提出合理的原子编辑,从而实现零阶优化策略。不过,零阶搜索存在“最弱链效应”:当一组编辑被整体接受或拒绝时,单个有害编辑会抵消所有其余编辑的益处。为解决该问题,我们引入HERO,这是一种程序优化器,它提示LLM生成多样、非重叠的原子编辑,随后使用评估器分数系统地选择并将这些编辑组合成连贯的程序改进。我们在算法问题、策略游戏、基于LLM的智能体系统设计以及机器人路径规划等领域对HERO进行评估,结果显示,在所有这些领域中,HERO均能持续发现分数更高的程序,收敛速度明显快于现有LLM优化器,同时消耗更少的token。

英文摘要

Large language models (LLMs) are increasingly used to solve complex problems by searching over program space, offering a general paradigm for scientific problems that can be naturally represented and solved as programs. Despite recent progress, identifying effective optimization directions for a candidate program remains challenging. By analogy with automatic differentiation, existing methods typically guide the search using a textual ``gradient'': a first-order update direction expressed as textual edits. Such gradients are inferred either from previously evaluated programs or from LLM-generated feedback on the implicit program-score mapping. However, these estimates become increasingly unreliable as the program--score mapping grows more complex, limiting their practical utility. We argue that explicit gradients are not essential for effective program optimization. Leveraging their prior knowledge, LLMs can propose plausible atomic edits directly from the current program, thereby enabling a zeroth-order optimization strategy. However, zeroth-order search suffers from a \textit{weakest-link effect}: when a bundle of edits is accepted or rejected as a whole, a single harmful edit can negate the benefits of all remaining edits. To address this issue, we introduce HERO, a program optimizer that prompts an LLM to generate diverse, non-overlapping atomic edits and then systematically selects and composes them into coherent program improvements using evaluator scores. We evaluate HERO across algorithmic problems, strategy games, the design of LLM-based agentic systems, and robotic path planning. Across these domains, HERO consistently discovers higher-scoring programs and converges substantially faster than prior LLM-based optimizers, while consuming fewer tokens.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑