arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

编码代理可以复现科学机器学习论文

Coding-agents can replicate scientific machine learning papers

Atharva Hans, Ilias Bilionis

arXiv 2607.02134首次发表:更新:

发表机构

School of Mechanical Engineering, Purdue University(机械工程学院,普渡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出Paper-replication工作流,让编码代理以目标驱动方式复现论文的计算声明,通过记录目标、重建方法、运行实验并验证,在四篇论文的12次独立运行中全部完成并通过验证。

AI 中文摘要

科学机器学习论文通常包含计算声明,例如相对均方误差小于5%或95%预测可信区间覆盖测试数据。编码代理可以被提示仅从论文材料复现这些声明,但提示本身不能可靠地保存进展或检查生成的证据是否支持论文的声明。我们引入了Paper-replication,这是一个工作流,使每个选定的论文声明成为带有记录证据的目标,并将其实现为编码代理技能。该工作流使代理记录这些目标,重建论文的方法,运行计算实验,将生成的输出与来源和论文声明的比较联系起来,记录匹配证据在复现报告中的位置,并在完成前通过验证检查。我们在四篇科学机器学习论文的十二次独立运行中评估了Paper-replication。所有十二个工作空间都通过了完成门,所有158个记录的目标都与报告覆盖范围匹配。即使在这种完成的工作空间状态下,重复运行在论文如何划分为目标、对源论文的数值保真度、经过的复现时间、在最终证据被接受之前替换的中间执行次数以及用于接受证据的规则方面存在差异。Paper-replication使完成依赖于工作空间证据和验证检查,而不是代理的最终消息。

英文摘要

Scientific machine learning papers typically make computational claims, e.g., that the relative mean square error is less than 5% or that the 95% predictive credible interval covers the test data. A coding agent can be prompted to replicate those claims from paper materials alone, but the prompt does not by itself reliably preserve progress or check whether generated evidence supports the paper's claims. We introduce Paper-replication, a workflow that makes each selected paper claim a target with recorded evidence, and implement it as a coding-agent skill. The workflow makes the agent record those targets, reconstruct the paper's method, run computational experiments, link generated outputs to provenance and comparisons with the paper's claims, record where matched evidence appears in the replication report, and pass validation checks before completion. We evaluate Paper-replication on twelve independent runs across four scientific machine learning papers. All twelve workspaces pass the completion gate, and all 158 recorded targets are matched with report coverage. Even in this completed workspace state, repeated runs differ in how papers are divided into targets, in numerical fidelity to the source papers, in elapsed replication time, in the number of intermediate executions replaced before final evidence is accepted, and in the rules used to accept evidence. Paper-replication makes completion depend on workspace evidence and validation checks rather than on the agent's final message.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑