arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33180stat.MLcs.AIcs.LGstat.ME

我们应该信任哪些自我改进?智能体复用基准时的可靠自我改进

Which Self-Improvements Should We Trust? Reliable Self-Improvement When Agents Reuse Their Benchmarks

  • Purdue University(普渡大学)

机构由 AI 辅助整理,请以论文原文为准。

Xiaojing Sun, Yuhan Zeng, Zihua She, Xiao Wang

中文总结 AI 辅助

针对递归自我改进中固定基准复用导致的自适应过拟合问题,提出 REUSE 框架,通过限制评估反馈和统计保证,将虚假提升比例从 20.7% 降至 0%,确保提升的真实性。

中文摘要 AI 辅助

随着递归自我改进(RSI)的快速发展,可靠的评估对于指导自适应搜索变得至关重要。RSI 通常依赖有限的评估资源(如固定基准)来决定保留哪些修改以及下一步提出什么。然而,当这些有限资源被反复复用时,新候选是基于同一评估集的反馈提出的,因此搜索轨迹可能自适应过拟合,经验改进可能无法反映底层任务分布上的真实群体改进。一些现有方法考虑了多重比较,但假设候选独立于评估集选择,因此不能控制这种自适应依赖性。为解决此问题,我们提出 REUSE(顺序演化下的风险控制评估),一个认证评估与提升框架,允许固定评估集支持重复的自适应决策,同时提供统计保证。对于用户指定的错误水平 $\alpha$,以至少 $1-\alpha$ 的概率,每个被提升的修改都是底层任务分布上的真实群体改进。REUSE 通过严格限制返回给搜索过程的评估反馈,并在错误预算内考虑可能的提升历史来实现这一点。我们为该设置下的 RSI 评估发展了详细的统计理论,包括同时错误控制、累积改进的有效下界,以及自适应评估复用的基本限制的特征化。在实时自我改进实验中,与当前 RSI 系统和错误控制基线中的评估框架相比,REUSE 显著减少了虚假提升,将虚假提升比例从最高 20.7% 降至 0%,同时实现了与最佳基线相当的真实群体性能。

英文摘要

As recursive self-improvement (RSI) rapidly advances, reliable evaluation becomes critical for guiding adaptive search. RSI typically relies on finite evaluation resources, such as fixed benchmarks, to determine which modifications are retained and what is proposed next. However, when these finite resources are repeatedly reused, new candidates are proposed based on feedback from the same evaluation set, so the search trajectory can adaptively overfit and empirical improvement may not reflect genuine population improvement on the underlying task distribution. Some existing methods account for multiple comparisons but assume that candidates are chosen independently of the evaluation set, and therefore do not control this adaptive dependence. To address this, we propose REUSE (Risk-controlled Evaluation Under Sequential Evolution), a certified evaluation and promotion framework that allows a fixed evaluation set to support repeated adaptive decisions while providing statistical guarantees. For a user-specified error level $α$, with probability at least $1-α$, every promoted modification is a genuine population improvement on the underlying task distribution. REUSE achieves this by strictly limiting the evaluation feedback returned to the search process and accounting for possible promotion histories within the error budget. We develop detailed statistical theory for RSI evaluation in this setting, including simultaneous error control, valid lower bounds on cumulative improvement, and a characterization of the fundamental limits of adaptive evaluation reuse. In live self-improvement experiments, REUSE commits substantially fewer false promotions than evaluation frameworks from current RSI systems and error-controlled baselines, reducing the proportion of false promotions from up to 20.7% to 0%, while achieving final true population performance comparable to the best baselines.

↑