发表机构
Stanford University; University of Wisconsin–Madison(斯坦福大学; 威斯康星大学麦迪逊分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究构建了含425项多领域假设检验任务的P-Bench基准,训练出Fisher-R1智能体,其在P-Bench上较DeepSeek-V4-Pro单试验成功率平均提升21%,证明强化学习可提升LLM的统计推理可靠性。
AI 中文摘要
可靠的假设检验是许多实证科学结论的基础。大语言模型(LLM)智能体越来越多地被用于自动化这一过程,因为它们可以检查数据集、生成代码并端到端地生成分析结果。然而,我们发现尽管这些智能体的分析执行正确,却经常出现细微的推理错误,导致结论错误。现有的基准测试未能捕捉到这种失败模式,因为它们很少评估报告的p值在给定数据潜在假设下是否具有统计有效性。我们通过构建P-Bench来解决这一缺口,该基准包含425个开放式、现实的假设检验任务,涵盖经济学、生物学和医学领域。每个任务要求智能体仅根据科学假设和数据集,选择统计方法、计算p值并得出结论。我们进一步引入Fisher-R1,这是一种使用合成任务和强化学习训练的、用于严格假设检验的开放权重LLM智能体。在P-Bench上,Fisher-R1-14B显著优于其骨干模型,也优于包括GPT-5.4和DeepSeekV4-Pro在内的强大专有和开源基准模型,与DeepSeek-V4-Pro相比,单试验成功率平均相对提升21%,在最具挑战性的任务上提升高达26%。我们的结果表明,当前LLM智能体缺乏用于假设检验的可靠统计推理能力,而在经过验证的统计奖励任务上进行强化学习可大幅提升其可靠性。
英文摘要
Reliable hypothesis testing is the foundation of many empirical scientific claims. Large language model (LLM) agents are increasingly used to automate this process, as they can inspect datasets, generate code, and produce analyses end-to-end. However, we show that they frequently make subtle inferential errors that lead to incorrect conclusions despite correctly executed analyses. Existing benchmarks fail to capture this failure mode, as they rarely assess whether a reported p-value is statistically valid given the assumptions underlying the data. We address this gap by building P-Bench, a benchmark comprising 425 open-ended, realistic hypothesis-testing tasks spanning economics, biology, and medicine. Each task requires an agent to select a statistical method, compute a p-value, and draw a conclusion given only a scientific hypothesis and a dataset. We further introduce Fisher-R1, an open-weight LLM agent trained for rigorous hypothesis testing using synthetic tasks and reinforcement learning. On P-Bench, Fisher-R1-14B substantially improves over its backbone and outperforms strong proprietary and open-source baselines, including GPT-5.4 and DeepSeekV4-Pro, achieving a 21% average relative improvement in single-trial success over DeepSeek-V4-Pro, with gains up to 26% on the most challenging tasks. Our results demonstrate that current LLM agents lack reliable statistical reasoning for hypothesis testing and that reinforcement learning on tasks with verified statistical reward substantially improves reliability.