BONSAI:基于可进化性的技能树搜索
BONSAI: Evolvability-Guided Tree Search over Skills
浏览论文内容
中文总结 AI 辅助
该研究提出BONSAI技能优化框架,基于可进化性的蒙特卡洛树搜索优化冻结智能体的技能,在30B智能体和三个基准上,较无技能智能体及GEPA、SkillOpt基线显著提升保留集准确率。
中文摘要 AI 辅助
技能是一种自然语言文档,用于指导权重无法更新的冻结智能体,因此智能体缺失的任何能力都必须以散文形式提供。优化技能本质上是针对分数优化文本,而保留任何能提高保留分数的编辑的标准方法存在特定缺陷:单个分数无法区分处于狭窄过拟合峰值的文档和处于宽阔平台的文档,尽管只有后者仍可改进。我们引入BONSAI,这是一种新颖的技能优化框架,其指导依据是可进化性,即文档空间区域在进一步变异下持续产生可行变异的能力——生物学将此属性与当前适应性视为不同概念。BONSAI将技能构建为蒙特卡洛搜索树,其中每个子文档是其父文档的变异体,并根据上置信度选择规则下降,该规则的利用项融合了技能自身的适应性及其变异邻域的适应性。由于每个子文档都是变异体,节点下记录的平均分数可估计该邻域的可进化性,且无需额外成本,因此该规则将预算集中在持续改进的区域,同时其探索项使当前较弱的分支保持竞争状态。BONSAI仅输出其找到的单份最佳分数文档,除了替换它的「接受即改进」循环外,无其他成本。在使用冻结的30B智能体并对三个基准进行平均后,BONSAI将保留集准确率比无技能智能体提高了23.13个百分点,且在两个预算匹配基线GEPA和SkillOpt上分别提高了3.87和3.97个百分点。
英文摘要
A skill is a naturallanguage document that steers a frozen agent whose weights cannot be updated so any capability the agent lacks must be supplied in prose Optimising a skill is therefore optimising text against a score and the standard recipe which keeps any edit that raises a heldout score is blind in a specific way a single score cannot tell a document perched on a narrow overfit spike from one resting on a broad plateau even though only the second can still be improved We introduce BONSAI a novel skilloptimisation framework that steers instead by evolvability the capacity of a region of documentspace to keep producing viable variation under further mutation a property biology treats as separate from present fitness BONSAI grows skills as a MonteCarlo search tree in which every child document is a mutation of its parent and descends it under an upperconfidence selection rule whose exploitation term blends a skills own fitness with the fitness of its mutational neighbourhood Because every child is a mutation the mean score recorded beneath a node estimates that neighbourhoods evolvability at no extra cost so the rule concentrates budget on regions that keep improving while its exploration term keeps a currently weak branch in contention BONSAI ships the single bestscoring document it finds at no cost beyond the acceptifbetter loop it replaces With a frozen 30B agent and averaged over three benchmarks BONSAI lifts heldout accuracy over the skillfree agent by 2313 points and improves on two budgetmatched baselines GEPA and SkillOpt by 387 and 397 points respectively