什么能证明你错:对可证伪的研究构思进行基准测试的语言模型
What Proves You Wrong: Benchmarking Language Models on Falsifiable Research Ideation
浏览论文内容
中文总结 AI 辅助
该研究提出 Lit2Test 基准,通过六领域契约让研究构想可证伪,对四个前沿模型的研究构想进行盲态成对比较,恢复了模型的严格排名并公开相关资源。
中文摘要 AI 辅助
大型语言模型正越来越多地被用于提出研究构想,然而评判这类构想的主流方式缺乏统一决策规则:自由形式的评判会受风格和立场影响,而针对后续论文的评分则是奖励对已实现轨迹的恢复。我们引入了一个将构想从文献推向测试的基准——Lit2Test 基准,它围绕可证伪结果组织成一个六领域契约,使得每个构想预先承诺了能证明其错误的观测结果,从而让其质量首先是可判定的,而非仅仅是可争论的。该基准前瞻性地基于 200 个真实论文邻域构建,从四个前沿模型中引出构想,并通过 1200 次成对比较进行评估,评估在两种呈现顺序下均为盲态。该方案通过诊断控制和有限的人类校准来审计自身可靠性,三名注释者在明确的可靠性范围内证实了结论。Lit2Test 在所有 10000 次自助抽样重复中都恢复了四个模型的严格排名,这种区分源于所提出测试和指标的质量,而非表面流畅性。我们发布该基准、构建流程和审计产物供公众使用。
英文摘要
Large language models are increasingly used to propose research ideas, yet the prevailing ways of judging such ideas supply no shared decision rule: free-form judging sways with style and position, and scoring against a later paper rewards recovery of one realized trajectory. We introduce a benchmark that carries a proposal from Literature to Test: the Lit2Test benchmark centers on a six-field contract organized around a falsifying outcome, so that every proposal precommits the observation that would prove it wrong, making its quality decidable in the first place rather than merely arguable. Built prospectively from 200 real-paper neighborhoods, Lit2Test elicits proposals from four frontier models and compares them through 1,200 pairwise comparisons judged blind in both presentation orders. The protocol audits its own reliability through diagnostic controls and bounded human calibration, with three annotators corroborating the conclusions within explicitly stated reliability bounds. Lit2Test recovers a strict ranking of the four models in all 10,000 bootstrap replicates, and the separation comes from the quality of the proposed tests and metrics rather than from surface fluency. We release the benchmark, construction pipeline, and audit artifacts for public use.