发表机构
Meta; Microsoft(Meta; 微软)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究揭示 LLM 自我改进循环中因小评估集导致的赢家诅咒,提出选择噪声模型,并实证表明保留增益常被高估,建议报告带不确定性的保留增益。
AI 中文摘要
自我改进的 LLM 系统会对自己提出修改,并保留那些在小规模评估集上得分更高的修改。我们将这种“若更好则保留”的步骤视为测量噪声下的选择,在单一决策中对候选者的相关误差进行建模,并实证研究当评估集被重复使用时会发生什么。在 Qwen 模型重写自身指令且每个候选者也在 600 个保留项目上评分的运行中,第一次之后的多数提案是有害的,模型给出了某代最佳候选者的赢家诅咒的大小。利用来自独立试点的先验,它匹配了原生循环中第一代提交的平均高估,尽管并非逐设置匹配。在一项预注册研究中,贪心循环的最终选择集得分超过保留准确率 13 到 20 个百分点(16 个选择项)和 1 到 5 个百分点(256 个选择项)。在 TREC 上,保留增益随选择集增大而增加,但在 GSM8K 上并非如此,并且所测试的接受规则在整个运行中并未优于贪心接受。当当前模型改进一条有能力的指令时,在选择集上测量的增益也超过了保留增益,并且在 GEPA 和 MIPROv2 的验证分数中也是如此。对起始指令和当前指令在 64 个从未用于选择的项目上进行评分,消除了循环报告增益的平均偏差,但单个估计仍偏差约 6 个百分点。自我改进研究应报告带有不确定性的保留增益。
英文摘要
Self-improving LLM systems propose changes to themselves and keep those that score better on a small evaluation set. We treat this keep-if-better step as selection under measurement noise, model the correlated errors of the candidates in a single decision, and study empirically what happens when the evaluation set is reused. In runs where Qwen models rewrite their own instructions and every candidate is also scored on 600 held-out items, most proposals after the first are harmful, and the model gives the size of the winner's curse of a generation's best candidate. With a prior from a separate pilot, it matches the average overstatement of first-generation commits in native loops, though not setting by setting. In a pre-registered study, the final selection-set score of greedy loops exceeded held-out accuracy by 13 to 20 points with 16 selection items and by 1 to 5 points with 256. Held-out gains grew with the selection set on TREC but not on GSM8K, and the tested acceptance rules did not beat greedy acceptance over whole runs. Gains measured on the selection set also exceeded held-out gains when a current model refined a competent instruction, and in the validation scores of GEPA and MIPROv2. Scoring the starting and the current instruction on 64 items never used for selection removes the average bias of a loop's reported gain, but single estimates remain off by about 6 points. Self-improvement studies should report held-out gains with their uncertainty.
Comments34 pages, 4 figures, 18 tables; code and saved experimental records included as ancillary files