arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.17590cs.NEcs.AI

进化集成搜索:具有持久记忆的委员会引导程序进化

Evolutionary Ensemble Search: Council-Guided Program Evolution with Persistent Memory

发表机构脉冲人工智能
查看机构详情
  • Impulse AI(脉冲人工智能)

机构由 AI 辅助整理,请以论文原文为准。

Juan P. Madrigal-Cianci, Eshan Chordia

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出进化集成搜索(EES)框架,通过委员会引导的程序进化与持久记忆,在MLE-bench Lite上19/22任务达奖牌阈值,实现跨模态机器学习流程的累积式自动构建。

中文摘要 AI 辅助

进化集成搜索(EES)通过专家引导的程序进化来构建机器学习流程。一个角色专业化的委员会将任务证据和实验结果转化为结构化的搜索方向。一个编排器将这些方向分配给执行专家和一个进化引擎。该引擎选择经过评估的父代,诊断其错误,并通过代码变异、结构化流水线编辑和交叉产生后代。每个子代必须执行并获得自己的验证证据。种群档案保留有用的备选方案,而兼容的预测在验证门控的集成阶段进行竞争。搜索通过父代相对算子信用、会话记忆和跨运行检索的问题索引课程进行自适应。我们详细说明了这些机制,区分了它们的执行概况,并定义了在候选计算发生变化时进行比较所需的契约。一个公开的MLE-bench Lite开发记录账本记录了22个任务中19个(86.36%)的奖牌阈值工件,最佳结果为11金、5银和3铜。这些流程涵盖文本、图像、表格、音频、科学几何和确定性变换。该活动包括运行之间的等级反馈、外部来源路径和混合确认程序;其总量是一个已实现的开发结果,而非盲目的自主智能体成功率。本报告为累积可执行搜索提供了一个具体架构,并对其跨模态开发结果进行了版本化记录。

英文摘要

Evolutionary Ensemble Search (EES) constructs machine-learning procedures through expert-guided program evolution. A role-specialized council turns task evidence and experimental results into structured search directions. An orchestrator allocates these directions to execution specialists and an evolutionary engine. The engine selects measured parents, diagnoses their errors, and produces descendants through code mutation, structured pipeline edits, and crossover. Each child must execute and acquire its own validation evidence. Population archives retain useful alternatives, while compatible predictions compete in a validation-gated ensemble stage. Search adapts through parent-relative operator credit, session memory, and problem-indexed lessons retrieved across runs. We specify these mechanisms, distinguish their execution profiles, and define the contracts required to compare candidates as their computations change. A public MLE-bench Lite development ledger records medal-threshold artifacts on 19 of 22 tasks (86.36\%), with best outcomes of 11 gold, five silver, and three bronze. The procedures span text, images, tables, audio, scientific geometry, and deterministic transformations. The campaign includes grade feedback between runs, external-source routes, and mixed confirmation procedures; its aggregate is an achieved development result, not a blind autonomous-agent success rate. The report contributes a concrete architecture for cumulative executable search and a versioned account of its cross-modal development outcomes.

↑