arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06817stat.AP

估计的潜在分布偏离正态性的情况有多常见?来自504个项目反应数据集的证据

How Common Are Estimated Latent-Distribution Departures From Normality? Evidence From 504 Item-Response Data Sets

JoonHo Lee

首次发表
浏览论文内容

中文总结 AI 辅助

该研究分析504个项目反应数据集,发现超半数的潜在分布与正态性差异显著,不同领域差异程度不同,灵活方法对异常数据集识别一致但幅度判断有分歧,建议应用分析明确分布假设并报告敏感性。

中文摘要 AI 辅助

项目反应模型通常假设特质分布为正态分布,但人们对实际研究中拟合的分布与正态性存在显著差异的频率,以及哪些报告结果受影响最大知之甚少。我们分析了来自项目反应仓库(Item Response Warehouse)中273项研究的504个项目反应数据集,对每个数据集分别拟合正态假设模型和从反应中估计的灵活分布模型,并比较两种校准下的信度、项目参数估计值、预测的测验反应和被试得分。在超过一半的数据集中,两种拟合分布在特质量表的某一点上的累积概率差异至少达到10个百分点,中位数最大差异约为11个百分点;约三分之一的数据集差异达到15个百分点,近五分之一的数据集差异达到20个百分点。估计的分布形状包括偏度、厚尾、平坦区域,偶尔还会出现多峰性。在多个态度、情感和行为领域,差异更大,不过当在项目模型族内比较数据集时,这些差异会减弱。当项目参数估计值保持固定时,信度通常变化很小,而重新拟合完整模型在某些数据集中会产生更大的变化,且项目参数估计值、预测的测验反应和被试得分的变化并不平行。不同的灵活方法通常会识别出相同的最异常数据集,但在差异幅度上存在一定分歧。应用分析应明确说明分布假设,检验灵活替代方案,并针对解释或决策中使用的每个结果分别报告敏感性。

英文摘要

Item response models usually assume a normal trait distribution, yet little is known about how often fitted distributions in real studies differ substantially from normality or which reported results are most affected. We analyzed 504 itemresponse data sets from 273 studies in the Item Response Warehouse, fitting each with the normal assumption and with a flexible distribution estimated from the responses, and we compared reliability, item estimates, predicted test responses, and person scores between the two calibrations. In more than half of the data sets the two fitted distributions differed by at least 10 percentage points of cumulative probability at some point on the trait scale, with a median maximum difference of about 11 points; about one-third reached 15 points, and nearly one-fifth reached 20 points. The estimated shapes included skewness, heavy tails, flat regions, and occasional multimodality. Differences were larger in several attitudinal, affective, and behavioral domains, although these contrasts weakened when data sets were compared within item-model families. Reliability usually changed little when item estimates were held fixed, whereas refitting the full model produced larger changes in some data sets, and item estimates, predicted test responses, and person scores did not change in parallel. Alternative flexible methods generally identified the same data sets as most unusual but disagreed somewhat about magnitude. Applied analyses should state the distributional assumption, examine a flexible alternative, and report sensitivity separately for each result used in interpretation or decision making.

↑