arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

网络心理测量学与大语言模型的结合:应用于信度审计的伊辛嵌入模型

Integrating Network Psychometrics and LLMs: The Ising-Embeddings-Model applied to Reliability Auditing

Matthias von Davier

arXiv 2608.26790首次发表:更新:

AI 中文总结

该研究结合网络心理测量学与LiRA,提出改进的伊辛嵌入模型,用于国际评估的信度审计,可利用全数据集完成评分一致性估计等任务。

AI 中文摘要

大规模评估中建构反应题的评分一致性通常通过双评分法估计,该方法使用小样本并假设反应间相互独立。我们提出一种结合网络心理测量学与语言整合信度审计(LiRA)的整合框架,采用改进的伊辛模型。该模型在二元正确标签上定义联合分布,其中成对交互项设为句子嵌入的余弦相似度,并加入项目难度的全局偏差参数。LiRA在语义邻域上的加权多数投票被证明可近似该伊辛模型的条件逻辑分布。该参数框架支持基准分数生成、不确定性量化及缺失标签插补;参数通过最大伪似然估计。该方法使用完整数据集,无需大量双评分,考虑语义依赖,并提供评分者不一致的诊断。LiRA可扩展方法与概率图模型的结合,为PIRLS、PISA、TIMSS等国际评估提供全面的信度评估工具。

英文摘要

Scoring consistency for constructed-response items in large-scale assessments is typically estimated through double-scoring, which uses small samples and assumes independence among responses. We present an integrated framework combining network psychometrics with the Linguistic-Integrated Reliability Audit (LiRA) via a modified Ising model. The model defines a joint distribution over binary correctness labels with pairwise interactions set to the cosine similarity of sentence embeddings and a global bias parameter for item difficulty. LiRA's weighted majority voting over semantic neighborhoods is shown to approximate the conditional logistic distributions of this Ising model. The parametric framework supports benchmark score generation, uncertainty quantification, and missing label imputation; parameters are estimated by maximum pseudo-likelihood. The approach uses the full dataset without requiring extensive double-scoring, accounts for semantic dependencies, and provides diagnostics for rater inconsistencies. The integration of LiRA's scalable methodology with a probabilistic graphical model offers a comprehensive tool for reliability assessment in international assessments such as PIRLS, PISA, and TIMSS.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑