arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

循证科学问题发现:一种结合历史回溯测试的框架

Evidence-Based Scientific Question Discovery: A Framework with Historical Backtesting

Hui Mao

arXiv 2608.09968首次发表:更新:

AI 中文总结

该研究提出一种结合历史回溯测试的循证科学问题发现框架,在系外行星大气领域验证其可生成被后续文献关注的科学问题,为科学问题发现提供了新方法。

AI 中文摘要

当前AI系统被优化用于回答问题,而科学事业的瓶颈出现在更早的阶段——发现值得研究的问题。我们提出一种框架,该框架将可追溯、可复现、范围受控的研究语料转化为可排序、可证伪的研究问题:证据被表示为携带主张的来源;检测、分类并经人工裁决的跨论文张力;留存的信号被细化为问题,并通过两阶段协议排序,该协议将科学优先级与执行优先级分离。我们在系外行星大气领域实例化该框架,该领域独特结合了文献、结构化目录和空间望远镜档案。在历史回溯测试中,所有由2021年之前可用证据生成的问题,均被系统从未见过的2021至2026年文献实质性涉及:其中两个问题得到解答,包括一个其前提后来被科学界明确反驳的问题,而排名第一的问题被独立提出且仍未解决。这些结果表明,从证据张力中进行系统性问题发现,能够揭示后续科研人员投入精力的问题。

英文摘要

Current AI systems are optimized for answering questions; the scientific enterprise is bottlenecked earlier, at discovering the questions worth investigating. We present a framework that turns a traceable, reproducible, scope controlled research corpus into ranked, falsifiable research questions: evidence is represented as provenance carrying claims; cross paper tensions are detected, typed, and human adjudicated; surviving signals are refined into questions and ranked by a two stage protocol separating scientific priority from execution priority. We instantiate the framework on exoplanet atmospheres, a domain that uniquely combines literature, structured catalogs, and space telescope archives. In a historical backtest, all questions generated from evidence available before 2021 were substantively engaged by the 2021 to 2026 literature the sys?tem never saw: two were answered, including one whose premise the community later explicitly refuted and the top ranked question is independently posed and still open. These results sug?gest that systematic question discovery from evidence tensions surfaces the questions working scientists subsequently invest in.

Comments13 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑