重构:从预发表参考文献中恢复研究思路的盲基准
Reconstruction: A Blind Benchmark for Recovering Research Ideas from Pre-Publication Bibliographies
浏览论文内容
中文总结 AI 辅助
该研究提出Reconstruction盲基准,测试语言模型从预发表参考文献恢复论文思路的能力,发现多智能体流水线可将匹配率提升约2.4倍。
中文摘要 AI 辅助
当仅获取一篇已发表论文的预发表参考文献时,语言模型能否恢复该论文的真实研究思路?我们提出Reconstruction,这是一种盲式思路恢复基准,它会隐藏种子论文以及所有同期或后续文献,要求模型提出假设,由一个独立的大型语言模型评判将这些假设与保留的真实思路进行匹配。该基准采用严格的防泄露协议,包括时间引用截断、匿名参考文献ID以及固定的单篇论文参考文献,以防止种子思路在提示阶段泄露。在六个科学领域的643篇评估论文中,七个前沿模型仅达到了中等匹配率(约3%-15%)。随后,我们评估了一个仅使用参考文献的多智能体(前4名)流水线,该流水线结合了跨模型评审与对齐假设槽的瑞士锦标赛,无需外部网络搜索。跨模型评审加锦标赛选择使所有六个领域的匹配率提升至约23%-42%,相比最佳单模型基线提升了约2.4倍。本草案报告了该协议、防泄露设计以及当前结果,时间为arXiv时间戳。
英文摘要
Can a language model recover the true research idea of a published paper when given only that paper's pre-publication bibliography? We introduce Reconstruction, a blind idea-recovery benchmark that withholds the seed paper and all contemporaneous or future literature, and asks models to propose hypotheses that an independent large language model judge matches against the held-out ground-truth idea. A strict anti-leakage protocol-temporal citation cutoff, anonymous reference IDs, and frozen per-paper bibliographies, which prevents prompt-time leakage of the seed idea. Across six scientific domains and 643 evaluated papers, seven frontier models achieve only modest Match rates (approx. 3-15%). We then evaluate a reference-only multi-agent (top 4) pipeline that combines cross-model review with a Swiss tournament over aligned hypothesis slots, without external web search. Cross-model review plus tournament selection raises Match rates to approx. 23-42% across all six domains, which is an observed approx. 2.4x lift over the best single-model baseline. This draft reports the protocol, anti-leakage design, and current results as an arXiv timestamp.
发表机构
- Stanford University(斯坦福大学)
- Titan Holdings(泰坦控股公司)
- Prentis AI(普伦蒂斯人工智能公司)
机构由 AI 辅助整理,请以论文原文为准。