AI 中文总结
该研究对889个多类型项目反应数据集测量边缘可靠性,发现低可靠性普遍存在,数据集间差异真实且受定义影响大,建议可靠性报告明确定义并附不确定性度量及两类系数。
AI 中文摘要
几乎每一项心理学定量研究都会报告信度系数,因此该领域对作者选择发表的信度了解甚多,但对心理学实际产生的数据的可靠性却知之甚少,因为已发表的系数需经过计算和报告内容的筛选决策。为此,我们直接测量可靠性,采用相同的估计量和规则,对来自项目反应库(Item Response Warehouse,一个包含认知测试、临床筛查工具、人格量表和态度量表的公开项目反应数据集集合)的889个数据集进行分析,得出三项发现:低可靠性现象普遍存在,即便采用宽松的可靠性定义,仍有30%的数据集低于传统的0.80阈值;数据集间的差异是真实存在的而非统计误差,因为估计噪声仅占该差异的约1%;最后,研究结果取决于可靠性的定义,采用严格定义时,低于0.80的数据集占比升至52%,该差异足以改变对该领域的结论。我们得出结论:可靠性报告应说明所使用的定义,附上不确定性度量,并同时提供严格和宽松的系数。
英文摘要
Nearly every quantitative study in psychology reports a reliability coefficient, so the field knows a great deal about the reliability that authors choose to publish. It knows much less about the reliability of the data psychology actually produces, because published coefficients pass through decisions about what to compute and what to report. We therefore measure reliability directly, applying the same estimators under the same rules to 889 datasets from the Item Response Warehouse, a public collection of item-response data that spans cognitive tests, clinical screeners, personality inventories, and attitude scales. Three findings emerge. Low reliability is common: even under a lenient definition of reliability, 30% of datasets fall below the conventional .80 threshold. The variation across datasets is real rather than statistical, since estimation noise accounts for only about one percent of it. Finally, the answer depends on the definition itself: under a strict definition the share below .80 rises to 52%, a difference large enough to change what one concludes about the field. We conclude that a reliability report should say which definition it uses, attach a measure of uncertainty, and give a strict coefficient alongside a lenient one.