arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16795cs.CEcs.AI

用于科学问题发现的历史回溯测试:协议与天文学试点

Historical Backtesting for Scientific Question Discovery: A Protocol and Astronomy Pilot

Hui Mao

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出可证伪的历史回溯测试协议评估科学问题生成系统,发布天文学实例验证发现,优先证据结构生成优于仅LLM提示,还指出结果分类问题,发布前瞻性无干扰测试实例。

中文摘要 AI 辅助

当前,生成科学研究问题的系统通过专家评分、大语言模型(LLM)作为评判者的评分或精选案例研究进行评估,所有这些方法均具有主观性,且无法证伪。我们将历史回溯测试形式化为一种替代方案:系统从在历史截止点冻结的语料库中生成问题,这些问题在接触后续文献之前被冻结,随后,一个时间上隔离的未来语料库将确定每个问题是否被后续解答、部分解决、被独立提出或被忽略,以及其潜在前提是否得到支持或反驳。该协议与模型无关:任何能生成冻结问题的系统均可被评分。我们发布了可复现的天文学实例,包含时间隔离的语料库、冻结的问题、可审计的标签、四个参考基准及提交接口。研究得出两项发现:其一,优先证据结构的生成优于仅使用LLM的提示:在生成器分解结合四截止点压力测试(2010-2024年,798个经评判的问题)中,最后一个窗口晚于模型训练时间,仅使用LLM的生成表现出记忆相关性,缺乏特定远见,而完全不使用模型权重的生成器则能找到在各个时代其前提被未来反驳的问题;其二,一项七评判者一致性研究(两名盲法人类标注者、五个评判模型、90个项目)指出问题出在结果分类而非评判者:两名细心人类的Cohen's kappa值为0.17,所有评判模型与专业标注者的一致性相当或更高(0.17-0.26),前沿模型之间的一致性为0.60——通过模型间一致性验证LLM评判者会将其可靠性高估三倍。我们还发布了一个前瞻性实例:2026年8月17日冻结的200个问题,将于2027-2030年评分,使核心主张成为由时间本身评判的无干扰测试。

英文摘要

Systems that generate scientific research questions are evaluated today by expert scores, LLM-as-judge ratings, or curated case studies -- all subjective, none falsifiable. We formalize historical backtesting as an alternative: a system generates questions from a corpus frozen at a historical cutoff, the questions are frozen before any access to later literature, and a temporally isolated future corpus then determines whether each question was subsequently answered, partially addressed, independently posed, or ignored, and whether its underlying premise was supported or refuted. The protocol is model-agnostic: any system that emits frozen questions can be scored. We release reproducible astronomy instances with temporally isolated corpora, frozen questions, auditable labels, four reference baselines, and a submission interface. Two findings result. First, evidence-structure-first generation outperforms LLM-only prompting: across a generator decomposition crossed with a four-cutoff stress test (2010-2024, 798 judged questions) whose last window postdates model training, LLM-only generation shows memorized relevance without specific foresight, while a generator using no model weights at all finds questions whose premises the future refutes in every era. Second, a seven-rater agreement study (two blinded human annotators, five judge models, 90 items) indicts the outcome taxonomy rather than the judge: two careful humans agree at kappa = 0.17, every judge model agrees with the professional annotator as well or better (0.17-0.26), and frontier models agree with one another at 0.60 -- certifying an LLM judge by model-model agreement would have overstated its reliability threefold. A prospective instance -- 200 questions frozen 2026-08-17, scored 2027-2030 -- is released so the central claims become contamination-free tests that time itself will grade.

补充信息

↑