arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

迈向基于嵌入的心理测量学:基于上下文分数的评估项目语义结构建模

Toward Embedding-Based Psychometrics: Structural Modeling of Assessment-Item Semantics With Contextual Scores

Jinsong Chen, Shi-Ting Chen

arXiv 2609.31976首次发表:更新:

AI 中文总结

本研究通过部分指定两步因子分析TIMSS数学项目语义,发现七组结构,支持条件性语义表示,但限制直接诊断解释,并展望文本辅助响应校准应用。

AI 中文摘要

上下文分数通过评估项目与外部语料库中参考词的相似性来表示这些项目。我们使用部分指定的两步因子程序,检验了40个TIMSS数学评分单元的分数语义结构。对因子数量的搜索在特定构建下识别出一个稳定的七组结构。随后的比较一致地支持一个一般维度与组关联并存,尽管个别组成员关系对某些规范选择仍然敏感。项目示例区分了重复出现的、跨领域的、敏感的和强加的关联。更简单且无限制的参考词阐明了锚定表示的贡献与局限:它优于单一因子,但未达到最低的工作贝叶斯信息准则(BIC)。一个独立的响应基准比较了三种初始Q构建及其在高阶和饱和属性分布下的Hull-PVAF修订。在这些诊断模型中,BIC倾向于官方内容框架,而赤池信息准则(AIC)倾向于其直接的四因子扩展,但一个匹配的单维双参数逻辑模型在所有十二种条件下的AIC和BIC均更低。这些发现支持一种条件性语义表示,同时限制了直接诊断解释。我们讨论了基于学习的文本辅助响应校准作为一项前瞻性应用,这需要更大的校准项目库和独立评估。

英文摘要

Contextual scores represent assessment items through their similarities to reference words in an external corpus. We examine the semantic structure of scores for 40 TIMSS mathematics scored units using a partially specified two-step factor procedure. A search across factor counts identifies a persistent seven-group structure under the featured construction. Subsequent comparisons consistently favor a general dimension alongside group associations, although individual group memberships remain sensitive to some specification choices. Item examples distinguish recurring, cross-domain, sensitive, and imposed associations. Simpler and unrestricted references clarify the contribution and limits of the anchored representation: it improves on a single factor but does not achieve the lowest working Bayesian information criterion (BIC). A separate response benchmark compares three initial Q constructions and their Hull-PVAF revisions under higher-order and saturated attribute distributions. Among these diagnostic models, BIC favors the official content framework and the Akaike information criterion (AIC) favors its direct four-factor augmentation, but a matched unidimensional two-parameter logistic model has lower AIC and BIC than all twelve conditions. These findings support a conditional semantic representation while limiting direct diagnostic interpretation. We discuss learned text-assisted response calibration as a prospective application requiring a larger calibrated item bank and independent evaluation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑