arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.27687cs.AI

Rehearse:从自改进自动研究中的置信度悬崖后退

Rehearse: Stepping Back from the Confidence Cliff in Self-Improving Autoresearch

Jiazhen Ji, Shouhong Ding

首次发表
浏览论文内容

中文总结 AI 辅助

该研究发现自动研究中存在置信度悬崖问题,提出Rehearse方法通过聚焦过往尝试结果记忆提升判断准确率,在三个任务的4000次训练运行预算下改善了终点性能。

中文摘要 AI 辅助

自动研究(Autoresearch)通过提出修改方案、运行完整训练任务并保留提升指标的修改来改进机器学习代码。该循环的效率不仅取决于生成想法的能力,还取决于智能体在花费一次训练运行之前判断拟议修改是否可能有效的能力。我们研究了这种执行前判断的可靠性在自动研究轨迹过程中如何变化。在公开的AutoSOTA日志(Li等人,2026;清华FIB实验室,2026)中,有用修改的比例从最初两次迭代的70%下降到第6次及以后迭代的43%。在来自39项论文衍生AutoSOTA任务的296个相同基线修改对中,每个对包含一个提高指标的修改和一个未提高指标的修改,且测量结果被隐藏,给定候选理由但无先前尝试历史的LLM判断器在严格共识给出裁决的对上达到79.5%的准确率。然而,在全部366对基准中,这种能力在循环后期大幅减弱。随着成功修改的积累,选择性准确率(基于严格共识裁决的条件准确率)从82.8%降至56.9%,而判断器仍愿意做出裁决。我们将这种操作模式称为置信度悬崖(confidence cliff)。Rehearse将循环修改实现为自动研究循环的轻量技能:提出多个想法,在执行前对其进行比较,运行最有前景的那个,并借助对类似过去尝试和结果的聚焦记忆进行判断。这种聚焦的结果记忆将后期选择性准确率提升至83.5%。在三个循环中4000次预算训练运行的情况下,Rehearse在nanochat、图像分类和时间序列预测任务中,在相同训练运行预算下提升了终点性能。

英文摘要

Autoresearch improves machine-learning code by proposing changes, running full training jobs, and keeping changes that improve the metric. The efficiency of this loop depends not only on generating ideas, but also on the agent's ability to decide, before spending a training run, whether a proposed modification is likely to work. We study how the reliability of this pre-execution judgment changes over the course of an autoresearch trajectory. In public AutoSOTA logs (Li et al., 2026; Tsinghua FIB Lab, 2026), the fraction of helpful modifications falls from 70% in the first two iterations to 43% by iteration 6+. On 296 same-baseline modification pairs from 39 paper-derived AutoSOTA tasks, each containing one modification that improved the metric and one that did not, with measured outcomes hidden, an LLM judge given candidate rationales but no prior-attempt history reaches 79.5% accuracy on the pairs where strict consensus returns a verdict. On the full 366-pair benchmark, however, this ability weakens substantially late in the loop. As successful changes accumulate, selective accuracy - accuracy conditioned on a strict-consensus verdict - falls from 82.8% to 56.9%, while the judge remains willing to decide. We call this operational pattern the confidence cliff. Rehearse implements the loop change as a lightweight skill for autoresearch loops: propose several ideas, compare them before execution, run the most promising, and judge with a focused memory of similar past attempts and outcomes. This focused outcome memory raises late selective accuracy to 83.5%. Across 4,000 budgeted training runs over three loops, Rehearse improves the endpoint under the same training-run budget on nanochat, image classification, and time-series forecasting.

发表机构

  • Tencent(腾讯)

机构由 AI 辅助整理,请以论文原文为准。

↑