通过快速树搜索实现自我改进
Self Improvement via Fast Tree-search
浏览论文内容
中文总结 AI 辅助
针对编码智能体自我改进成本高的问题,提出SIFT框架,利用LLM评判信号和分解树搜索,在严格预算下显著提升编码性能并降低资源消耗。
中文摘要 AI 辅助
编码智能体可以递归地修改自身的实现,形成自我改进的循环。虽然先前的研究表明这可以提升编码基准测试的性能,但现有方法成本高昂且计算密集。我们引入了一个简单、样本高效的自我改进框架,在严格的预算约束下显著提升了编码性能。我们识别出候选自我修改的评估是主要的运行时瓶颈,因为先前的方法通过使用修改后的智能体重新运行一部分基准任务来估计其有效性,这非常耗时。我们提出了通过快速树搜索实现递归自我改进(SIFT),该方法用LLM作为评判者的信号来增强这些下游任务评估,该信号对候选补丁进行成对比较,胜负记录通过正则化的Bradley-Terry模型进行聚合,得到的强度分数驱动轻量级分解树搜索中的基于排名的父代采样。昂贵的下游任务评估仅保留给最有希望的节点。使用完全分解的树搜索流程,评判分数提供中间信号来指导对有希望的候选补丁的探索,而不会受到缓慢评估运行的瓶颈限制。SIFT在完整的Polyglot基准上优于现有的基于树搜索的自我进化框架,同时在CPU小时数、墙钟时间和API成本方面显著降低了资源需求。
英文摘要
Coding agents can recursively modify their own implementations, forming a loop of self-improvement. While prior work shows this can boost performance on coding benchmarks, existing approaches are costly and compute-intensive. We introduce a simple, sample-efficient self-improvement framework that significantly improves coding performance under strict budget constraints. We identify evaluation of candidate self-modifications as the main runtime bottleneck since prior approaches estimate their effectiveness by re-running a subset of benchmark tasks with the modified agent, which is time-consuming. We introduce Recursive Self Improvement via Fast Tree-search (SIFT), which augments these downstream task evaluations with an LLM-as-a-judge signal that performs pairwise comparisons between candidate patches, where the win-loss record is aggregated with a regularized Bradley-Terry model, and the resulting strength scores drive rank-based parent sampling inside a lightweight disaggregated tree search. Expensive downstream task evaluations are reserved only for the most promising nodes. Using a fully disaggregated tree search pipeline, the judge scores provide intermediate signal to guide exploration on promising candidate patches without being bottlenecked by slow evaluation runs. SIFT outperforms existing tree-search based self-evolution frameworks on the full Polyglot benchmark with significantly lower resource requirements in terms of CPU hours, wall clock time, and API cost.
发表机构
- Massachusetts Institute of Technology(麻省理工学院)
- Sakana AI
机构由 AI 辅助整理,请以论文原文为准。