arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

统计显著发现的可复制性如何?

How Replicable Are Statistically Significant Findings?

Patrick Vu, Stefan Faridani

arXiv 2608.23257首次发表:更新:

AI 中文总结

该研究分析不同领域已发表研究的预期重复概率,发现p值0.05的发现重复概率低,低可复制性源于原始研究低统计效力,大样本经济学文献重复概率仍低,指出单一研究的统计显著性仅为提示性证据。

AI 中文摘要

在实证科学中,显著性阈值通常决定发现是否被视为效应存在的证据。本文研究恰好达到常规显著性阈值的发现,在相同样本量的重复实验中仍保持显著的可能性有多大。为回答该问题,我们针对实验经济学、心理学和社会科学领域已发表研究,估计了给定p值条件下的预期重复概率。我们验证该指标能准确预测实际重复结果,表现优于预测市场。p值为0.05的发现,在不同领域的预期重复概率介于0.10至0.25之间。低可复制性反映了原始研究的低统计效力,而非发表偏倚。随后,我们开发了一种非参数估计量,并将其应用于使用更大样本的经济学文献,发现其重复概率更高但仍然较低。这些结果表明,单一研究中的统计显著性仅为效应存在的提示性证据,更有力的结论需要累积证据支持。

英文摘要

In the empirical sciences, significance thresholds often determine whether findings are treated as evidence of an effect. This paper studies how likely findings that just meet conventional significance thresholds are to remain significant in replications of the same sample size. To answer this question, we estimate the expected replication probability conditional on a given p-value among published studies for experimental economics, psychology, and social science. We validate this measure by showing it accurately predicts actual replication outcomes, outperforming prediction markets. A finding with a p-value of 0.05 has an expected replication probability ranging from 0.10 to 0.25 across fields. Low replicability reflects low power in original studies rather than publication bias. We then develop a nonparametric estimator and apply it to economics literatures that use larger samples, finding higher but still low replication probabilities. These results indicate that statistical significance in a single study provides only suggestive evidence of an effect. Stronger conclusions require cumulative evidence.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑