发表机构
University of Chinese Academy of Sciences; Institute of Software, Chinese Academy of Sciences; Alibaba Group(中国科学院大学; 中国科学院软件研究所; 阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究旨在评估大语言模型在开放式结论前阶段的发现能力,引入前瞻性假设发现任务及评估框架HypoArena,提出回顾性上下文回归构建基准数据,实验揭示模型能力分层及效应,结果支持将其作为评估大语言模型制定调查方向的独特目标。
AI 中文摘要
大语言模型在回答预先指定的问题方面表现出色,但其在开放式、结论前阶段的发现能力仍未得到充分衡量。我们引入了前瞻性假设发现(PHD),要求模型从不确定的证据(包括异常观测和碎片化记录)中自主构建有根据、有区分性且可测试的假设空间,以指导后续调查。为评估此能力,我们引入了HypoArena,包括HypoData(六个科学和分析领域的988个案例基准)和HypoEval(开放式假设集评估框架)。为大规模构建HypoData,我们提出回顾性上下文回归,这是一个通过去除明确结论、目标假设和回顾性因果归因同时保留事实基础,从完整专家文档重建结论前上下文的Forge - Audit管道。由于PHD允许多个有效输出,HypoEval结合双向成对判断与Bradley - Terry - Davidson聚合进行排名以及六维评分细则进行诊断。对15个前沿大语言模型的实验揭示了明显的能力分层和结构化分析技能的模型依赖效应,一些性能较低的模型在HypoArena上有所提升,而其他系统包括一个顶级模型出现倒退。与绝对评分细则评分相比,竞技场评估解决了模型之间更细粒度差异,聚合排名与人类专家和独立评判者高度一致。这些结果支持将PHD视为评估大语言模型在不给出最终结论时如何制定调查方向的一个独特目标。我们的代码和数据可在指定网址公开获取。
英文摘要
Large language models (LLMs) excel at answering pre-specified questions, yet their ability to navigate the open-ended, pre-conclusion stage of discovery remains largely unmeasured. We introduce Prospective Hypothesis Discovery (PHD), which asks models to autonomously construct grounded, discriminative, and testable hypothesis spaces from inconclusive evidence, including anomalous observations and fragmented records, to guide subsequent investigation. To evaluate this capability, we introduce HypoArena, comprising HypoData, a benchmark of 988 cases across six scientific and analytical domains, and HypoEval, an evaluation framework for open-ended hypothesis sets. To construct HypoData at scale, we propose Retrospective Context Regression, a Forge--Audit pipeline that reconstructs pre-conclusion contexts from completed expert documents by removing explicit conclusions, target hypotheses, and retrospective causal attributions while preserving the factual substrate. Because PHD admits multiple valid outputs, HypoEval combines bidirectional pairwise judgments with Bradley--Terry--Davidson aggregation for ranking and six-dimensional rubric scoring for diagnosis. Experiments on 15 frontier LLMs reveal clear capability stratification and model-dependent effects of structured analytical skills, with gains for several lower-performing models on HypoArena but regressions for other systems, including a top-performing model. Compared with absolute rubric scoring, arena evaluation resolves finer-grained differences among models, with aggregated rankings showing strong agreement with human experts and an independent judge. Together, these results support treating PHD as a distinct target for evaluating how LLMs formulate investigative directions when final conclusions are withheld. Our code and data are publicly available at github.com/SKYLENAGE-AI/HypoArena and github.com/SKYLENAGE-AI/HypoArena.