arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一次运行并非一个想法:自动化研究中的实施彩票

The Implementation Lottery: Auditing Idea Reliability in Automated Research

Jingjie Ning, Shanshan Zhong, Xiaochuan Li, Ji Zeng, Chenyan Xiong

arXiv 2607.26587首次发表:更新:

发表机构

Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对自动化研究中存在的“实施彩票”问题,提出想法可靠性审计方法,通过312次分配实验发现实施方差远大于成果重运行方差,表明需多实施证据支撑想法决策。

AI 中文摘要

自动化研究系统会利用实验分数来交付成果,并决定哪些想法需要保留、迁移和推进。然而,一次运行的分数仅对应一个想法的一次实施。将这种实现层面的分数视为关于其母机制的证据,便产生了“实施彩票”——即想法层面的结论取决于所抽样的是哪一种合理的实施方式。只要一次运行就能更新对某一机制的信念,这种不匹配就是结构性的。我们对其影响程度进行了估算。“想法可靠性审计”通过以下方式测量“想法可靠性”:验证并冻结候选卡片、对新会话的实施方式进行抽样、使用结果盲法的保真度标签,以及重新运行已保存的成果。它会报告想法的组内相关系数(ICC)和留一实施法(LOO)的胜者反转情况。以往研究通常重复任务;而我们重复想法。在13个表格任务和两种编码智能体设置下的312次分配中,实施方差分别是相同成果重运行方差的5倍和10倍以上,且在25.6%和43.6%的决策中,一次实施抽样的胜者与另外两次实施均值下的胜者不同。在两种结果盲法审查规则下,反转在卡片层面过滤后仍然存在。对三种具有确定性评估器的材料回归工作流的探索性诊断也发现,实施变异在分解中占主导地位。这些发现区分了想法可靠性与N个成果中的最优效用。在分数指导想法层面的分支、迁移或研究记忆之前,证据应涵盖多种实施方式。

英文摘要

Automated research agents use program scores to judge ideas. We call variation in this evidence across implementations the implementation lottery. We introduce an Idea Reliability Audit that freezes mechanism specifications, samples independent programs, and compares selected code with fresh implementations of its mechanism. Across 3,048 assignments on 31 tabular classification tasks, all four primary aggregation tests have Holm-adjusted $p\geq0.56$. Under mean-of-five selection, the prespecified secondary intention-to-treat comparison gives selected-code premiums of 0.38 [0.08, 0.82] and 0.45 [0.05, 1.12] accuracy-equivalent points for Bounded and Agentic execution, respectively. Minimum task-deletion means are 0.20 and 0.14. Fidelity conditioning exposes concentration: the Agentic premium falls from 0.33 to 0.01 when one task is removed. Post-outcome analysis finds cross-split variation on 41 of 70 paired cards per process. Exploratory replay gives nearly equal Bounded point losses at four and twenty implementations; Agentic point losses decrease across the four evaluated budget rules. Under duration costs, one seed per program minimizes fitted common-design variance. The two-seed design becomes preferable when implementation-to-seed cost ratios exceed approximately 13 or 9.5 in the continuous-budget model. The audit distinguishes evidence for reusing a selected artifact from evidence for implementing its idea again.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑