发表机构
Zhejiang Lab(之江实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出端到端复现框架评估LLM智能体重建隐性科学知识的能力,通过14项天文学研究(含13篇Nature论文)发现多数存在方法模糊性,且匹配结果不等于重建推理,瓶颈常在于连接相关信息而非检索。
AI 中文摘要
大型语言模型(LLM)融入科学工作流程的步伐正在加快,然而,它们重建已发表研究所依据的推理过程的能力仍未得到探索。论文明确规定了具体步骤,却将许多方法论依赖——数据选择、校准修正、先验和领域假设——留作隐性内容。这种模糊性使得对基于LLM的智能体的评估变得复杂,因为未能复现结果可能既反映了智能体的局限性,也可能反映了原始文献中的规定不足。我们提出了一个通过端到端复现来评估智能体的框架,将执行与验证分离,并将计算失败与方法论模糊性区分开来。我们将该框架应用于十四项天文学研究:一项来自《天体物理学报》的案例研究以及十三篇发表在《自然》期刊上的论文。其中十三篇中有十一篇包含阻碍唯一复现路径的模糊性。在一项受控案例研究中,十二条预定义路径,即针对样本定义、天空掩模和视差零点处理的三因素二水平敏感性分析(3x2x2),对同一量给出了从2.16到3.53千秒差距(kpc)的估计值,其中只有一条路径恢复了已发表的值(约2.70 kpc)。已发表的值从未被用作优化目标、选择标准或停止条件;匹配的路径是在所有十二条路径运行完毕后才发现。关键在于,决定性的信息(+0.02毫角秒的视差零点修正)已经存在于论文中,但智能体直到分析使该效应变得可见时才认识到其因果相关性。因此,匹配已发表的结果并不能验证对底层推理的重建,瓶颈往往在于未能连接相关信息,而非未能检索信息。因此,端到端复现既可作为可重复性的测试,也可作为评估AI系统中隐性科学知识的框架。
英文摘要
The integration of large language models (LLMs) into scientific workflows is accelerating, yet their ability to reconstruct the reasoning underlying published research remains unexplored. Papers specify explicit procedures while leaving many methodological dependencies-data selection, calibration corrections, priors, and domain assumptions-implicit. This ambiguity complicates the evaluation of LLM-based agents, since a failure to reproduce a result may reflect either limitations of the agent or underspecification in the source. We present a framework that evaluates agents through end-to-end reproduction, separating execution from verification and computational failure from methodological ambiguity. We apply it to fourteen astronomy studies: a case study from The Astrophysical Journal and thirteen papers published in Nature. Eleven of the thirteen contained an ambiguity preventing a uniquely specified reproduction path. In a controlled case study, twelve predefined paths, a 3x2x2 sensitivity analysis over sample definition, sky masking, and parallax zero-point treatment-gave estimates from 2.16 to 3.53 kpc for the same quantity, with only one recovering the published value (about 2.70 kpc). The published value was never used as an optimization target, selection criterion, or stopping condition; the matching path was found only after all twelve had run. Crucially, the decisive information (a +0.02 mas parallax zero-point correction) was already in the paper, but the agents did not recognize its causal relevance until the analysis made the effect visible. Matching a published outcome therefore does not validate reconstruction of the underlying reasoning, and the bottleneck is as often a failure to connect relevant information as to retrieve it. End-to-end reproduction thus serves both as a test of reproducibility and as a framework for evaluating implicit scientific knowledge in AI systems.