arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13590stat.MEstat.AP

连续演化IRT题库的贝叶斯共识校准

Bayesian Consensus Calibration of Continuously Evolving IRT Item Banks

Paul A. Jewsbury, Steven W. Nydick, Manqian Liao, Siyuan, Chen

首次发表
浏览论文内容

中文总结 AI 辅助

针对AI生成题库规模大、更新频繁导致传统IRT重拟合成本高的问题,提出共识校准方法,通过分治策略独立校准各时期并分层重建合并后验,在大规模评估中验证了其有效性。

中文摘要 AI 辅助

基于人工智能的题目生成和基于自然语言处理的题目参数预测,正在产生比传统题库规模更大、更稀疏且更新更频繁的题库。分层贝叶斯项目反应理论(IRT)是此类题库的自然校准框架,但每次更新时重新拟合整个累积作答历史的常见做法成本高昂,且可能超出可用内存。我们描述了共识校准(consensus calibration),这是一种分而治之的程序,它独立校准每个时间段,并通过两层结构重建合并后验。首先,每个时期的后验抽取通过稳健的特征曲线链接(Haebara)映射到共同度量,该链接对每次抽取分别求解,从而将链接变换的不确定性传播到链接后的后验中。其次,链接后的项目后验作为高斯密度的乘积进行组合,从中移除每个时期贡献的总体先验,并重新引入通过各时期总体后验共识获得的单一先验。该校正针对后验离散度,而不仅仅是其位置。作为共识校准的证据,我们在一项大规模操作化评估中,将其与合并单次运行分析在项目参数恢复、按暴露度不确定性诊断和能力分布方面进行了比较。

英文摘要

AI-based item generation and NLP-based prediction of item parameters are producing item banks that are substantially larger, sparser, and more frequently updated than conventional banks. Hierarchical Bayesian item response theory (IRT) is a natural calibration framework for such banks, but the common practice of refitting the entire accumulated response history at each update is costly and can exceed available memory. We describe \emph{consensus calibration}, a divide-and-conquer procedure that calibrates each time period independently and reconstructs the pooled posterior in two layers. First, the posterior draws of each period are mapped to a common metric by a robust characteristic-curve linking (Haebara) that is solved separately for each draw, which propagates the uncertainty of the linking transformation into the linked posteriors. Second, the linked item posteriors are combined as a product of Gaussian densities from which the population prior contributed by each period is removed and a single prior---obtained by consensus across the per-period population posteriors---is reinstated. The correction targets the posterior dispersion, not only its location. As evidence for consensus calibration, we compare it to a pooled single-run analysis on a large operational assessment in terms of item-parameter recovery, an uncertainty-by-exposure diagnostic, and the ability distributions.

发表机构

  • Duolingo(多邻国)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑