arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从搜索到信号:自动启发式设计中的在线后训练

From Search to Signal: Online Post-Training in Automatic Heuristic Design

Yilun Yuan, Tianyu Zhou, Zhenzhou Tang

arXiv 2609.39383首次发表:更新:

发表机构

Wenzhou University(温州大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出在自动启发式设计中,将在线后训练形式化为上下文相关信号构建,通过替代映射将程序有效性、任务性能和生成上下文转化为更新信号,以提升模型启发式设计能力。

AI 中文摘要

基于大型语言模型(LLM)的自动启发式设计(AHD)通过迭代地提出并改进启发式方法,将设计原理与可执行代码配对。任务特定的评估器对程序进行评估;执行结果和性能分数指导搜索过程。许多AHD系统保持生成器冻结不变;EvoTune和算法与语言模型协同进化(CALM)则根据评估后的候选方案更新生成器。当此类结果驱动带有可验证奖励的强化学习(RLVR)时,会形成一个搜索耦合循环:评估后的候选流既提供搜索状态更新,又为生成未来候选的模型提供训练信号。然而,有效性和性能并不能唯一确定有用的模型更新;将其转换为学习信号必须考虑生成每个候选方案时的提示和不断演化的搜索状态。我们将AHD中小型开放权重LLM的在线后训练形式化为上下文相关的信号构建,并开发了从程序有效性、任务性能和生成上下文到更新信号的替代映射。利用共享的评估轨迹和匹配的更新预算,跨AHD任务和模型家族的控制实验将这些映射与在线后训练基线进行比较,测试其对有效性、有效提案中的性能以及在上下文比较下改进的有效提案产出率的影响。互补的检查点、冻结搜索和实时系统评估检验了提案级增益是否体现在更新后的检查点行为和后续搜索中,而非仅来自累积的搜索状态。在预先指定的成本核算下进行的资源匹配比较,测试了在线更新是否在冻结生成器的额外搜索之外增加了价值。总之,这种设计避免了将端到端的搜索增益单独视为更强启发式设计能力的证据。

英文摘要

Large language model (LLM)-based automatic heuristic design (AHD) iteratively proposes and refines heuristics, pairing design rationales with executable code. Task-specific evaluators assess programs; execution outcomes and performance scores guide search. Many AHD systems keep the generator frozen; EvoTune and Co-Evolution of Algorithms and Language Model (CALM) instead update it from evaluated candidates. When such outcomes drive reinforcement learning with verifiable rewards (RLVR), they create a search-coupled loop: the evaluated candidate stream supplies both search-state updates and training signals for the model that generates future candidates. Yet validity and performance do not uniquely determine useful model updates; converting them into learning signals must account for the prompt and evolving search state that produced each candidate. We formulate online post-training of small open-weight LLMs in AHD as context-dependent signal construction and develop alternative mappings from program validity, task performance, and generation context to update signals. Using shared evaluated rollouts and matched update budgets, controlled experiments across AHD tasks and model families compare these mappings with online post-training baselines, testing their effects on validity, performance among valid proposals, and the yield of valid proposals that improve under contextual comparisons. Complementary checkpoint, frozen-search, and live-system evaluations assess whether proposal-level gains appear in updated checkpoint behavior and subsequent search, rather than arising solely from accumulated search state. A resource-matched comparison under pre-specified cost accounting tests whether online updating adds value beyond additional search with a frozen generator. Together, this design avoids treating end-to-end search gains alone as evidence of stronger heuristic-design capabilities.

Comments18 pages, including supplementary material. Preprint

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑