AI 中文总结
本研究提出一种基于载荷矩阵条件数的诊断指标,验证语义文本嵌入在多维自适应测试中可接近拟合参数的潜在特质恢复准确率,但存在后验方差膨胀问题,为替代项目校准提供评估依据。
AI 中文摘要
多维项目反应理论依赖于校准后的项目参数,如区分度和类别阈值,这些参数通常从大量人类测试响应样本中估计得出。本研究探讨这些参数的方向载荷是否可通过预训练句子嵌入直接从项目文本中恢复,从而避免初始项目校准的需求。我们使用开源IPIP大五人格数据集(样本量n=19719,含50个项目),采用D-最优项目选择构建了多维计算机化自适应测试(CAT)仿真。我们对比了三类项目载荷来源:拟合等级反应模型参数、语义文本嵌入和词汇基线。仿真结果显示,语义嵌入恢复潜在特质轮廓的准确率几乎与拟合参数相当(相关系数分别为0.825和0.857),明显优于简单词重叠方法(0.752)。然而,基于嵌入的模型产生了膨胀的后验方差,尽管点估计准确,但测量不确定性高出近4倍。我们将此归因于维度间的共线性,因为嵌入衍生的载荷在各特质上指向相似方向(条件数为137,对比1.0;平均特质余弦值为0.90)。这一结果反映了人格项目中常见的共享词汇。我们提出了一种基于载荷矩阵条件数的简单诊断指标,用于在测试前评估项目库是否适合采用文本衍生的载荷。
英文摘要
Multidimensional item response theory relies on calibrated item parameters, such as discrimination and category threshold values, which are usually estimated from large samples of human test responses. This study investigates whether the directional loadings of these parameters can be recovered directly from item text using pre-trained sentence embeddings, avoiding the need for initial item calibration. Using the open-source IPIP Big-Five dataset ($n=19{,}719$; 50 items), we built a multidimensional computerized adaptive testing (CAT) simulation using D-optimal item selection. We compared three item loading sources: fitted graded response model parameters, semantic text embeddings, and a lexical baseline. In simulation, semantic embeddings recovered latent trait profiles almost as accurately as fitted parameters (correlation $0.825$ vs. $0.857$), performing noticeably better than simple word overlap ($0.752$). However, the embedding-based model produced inflated posterior variance, showing nearly four times higher measurement uncertainty despite accurate point estimates. We attribute this to collinearity across dimensions, as embedding-derived loadings pointed in similar directions across traits (condition number $137$ vs. $1.0$; mean trait cosine $0.90$). This outcome reflects the shared vocabulary common in personality items. We propose a simple diagnostic metric based on the loading matrix condition number to evaluate whether an item bank is suitable for text-derived loadings prior to testing.