arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自演化算法设计智能体:通过种群策展策略优化逃离上下文演化停滞

Self-Evolving Algorithm-Design Agents: Escaping In-Context Evolutionary Stagnation via Population-Curated Policy Optimization

Chen Lu, Ke Xue, Siyuan Xu, Mingxuan Yuan, Chao Qian

arXiv 2609.38757首次发表:更新:

发表机构

State Key Laboratory of Novel Software Technology, Nanjing University; School of Artificial Intelligence, Nanjing University(南京大学计算机软件新技术全国重点实验室; 南京大学人工智能学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对算法设计智能体在上下文演化中易停滞的问题,提出种群策展策略优化(PCPO),通过全局种群与混合策略更新实现样本高效自演化,在芯片布局与GPU内核设计任务上超越现有方法。

AI 中文摘要

大语言模型正以算法设计智能体的形式越来越多地参与到复杂的现实世界任务中,负责设计和改进算法。许多成功的算法设计智能体采用纯上下文演化框架,但在需要专业知识的领域可能很快陷入平台期。参数化自适应提供了一种内化专业知识的方式,但传统训练需要丰富的领域特定语料库,而在复杂的算法设计场景中,高质量算法十分稀缺。在本文中,我们提出了样本高效的参数化自演化方法,使智能体能够探索并从自身生成的算法中学习。首先,我们刻画了上下文演化停滞的特征,并分析性地提出了改进链命题,展示了学习连续的自身生成算法如何局部增加邻近算法的似然性。受这种局部迁移视角的启发,我们进一步提出了种群策展策略优化(PCPO),利用全局种群和混合策略更新方案来保留和重用高质量、多样化的自身生成算法,将策略向更强的算法转移。在电子设计自动化中全局布局的学习率调度设计任务中,仅在4个芯片案例上训练,PCPO在16个芯片案例上的平均性能优于最先进的上下文演化方法(如OpenEvolve和ShinkaEvolve)。使用8B大小的基础模型,PCPO达到了与前沿闭源模型(如GPT-5.5)相当的性能。PCPO还通过内化基础领域知识和提示蒸馏降低了推理时的token成本。此外,PCPO在四个GPU内核设计上实现了显著的加速,相对于PyTorch Eager基线平均加速8.27倍。

英文摘要

Large language models are increasingly participating in complex real-world tasks in the form of algorithm-design agents, designing and refining algorithms. Many successful algorithm-design agents adopt pure in-context evolutionary frameworks, but they may quickly plateau in domains that require specialized knowledge. Parametric adaptation offers a way to internalize specialized knowledge, but conventional training requires abundant domain-specific corpora while high-quality algorithms are scarce in complex algorithm-design scenarios. In this paper, we propose sample-efficient parametric self-evolution where agents can explore and learn from self-generated algorithms. First, we characterize in-context evolutionary stagnation and analytically propose the Improvement Chain proposition, showing how learning successive self-generated algorithms can locally increase the likelihood of neighboring algorithms. Motivated by this local-transfer perspective, we further propose Population-Curated Policy Optimization (PCPO) to utilize a global population and a hybrid policy update scheme for retaining and reusing high-quality, diverse self-generated algorithms, shifting the policy towards stronger algorithms. In the task of learning rate schedule design for global placement in electronic design automation, trained only on 4 chip cases, PCPO outperforms the state-of-the-art in-context evolutionary methods (e.g., OpenEvolve and ShinkaEvolve) on average across 16 chip cases. With an 8B-size base model, PCPO achieves competitive performance compared to frontier closed-source models such as GPT-5.5. PCPO also reduces inference-time token cost by internalizing grounded domain knowledge and prompt distillation. Moreover, PCPO achieves significant speedups on four GPU kernel designs, with an average of 8.27$\times$ speedup against the PyTorch Eager baseline.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑