arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

对齐大语言模型模拟与人类应试者的心理测量校准:一种认知诊断画像方法

Aligning LLM-Simulated and Human Examinees for Psychometric Calibration: A Cognitive Diagnostic Profiling Approach

Wenjie Zhou, Yunting Liu, Renjiao Tang, Mark Wilson

arXiv 2607.26317首次发表:更新:

发表机构

University of California, Berkeley; University of Minnesota Twin Cities(加州大学伯克利分校; 明尼苏达大学双城分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出认知诊断画像(CDP)零样本框架,优化LLM模拟应试者的心理测量对齐,在Tatsuoka分数减法数据集上提升了能力分布、掌握画像及题目难度与人类的匹配度,助力LLM模拟应试者用于测试开发。

AI 中文摘要

教育测试的心理测量校准通常需要成本高昂的人类作答数据,而大语言模型(LLM)模拟应试者为早期校准提供了有前景的途径,但其作答过于准确且过于统一。我们提出认知诊断画像(CDP),这是一种零样本框架,通过提示LLM模拟具有多样化认知画像的合理应试者:将二元属性掌握模式呈现为自然语言画像,并在无信息或有信息分布下进行采样。使用Tatsuoka分数减法数据集(536名应试者、15个题目、5个属性),我们在无画像、无信息CDP和有信息CDP条件下评估了8种LLM配置,在能力分布、掌握画像和题目难度层面评估与人类应试者的对齐情况。CDP在所有三个层面均实现了提升:跨配置的分布重叠度上升;画像层面分数与人类画像预期的加权相关系数达到0.92至0.98;题目难度的恢复在秩次和绝对对齐上均有改善,对于具备推理能力的模型提升最为显著;在最优案例中,Gemini 3.0 Flash(Thinking)的单参数逻辑斯蒂(1PL)难度斯皮尔曼相关系数从0.24分别升至0.86和0.90,均方根误差(RMSE)从6.31降至1.30和0.90;有信息条件在画像层面对齐较强时帮助最大。CDP使LLM模拟应试者与人类应试者的心理测量对齐更紧密,使其可实际用于业务测试开发。

英文摘要

Psychometric calibration for educational tests typically requires costly human response data. Large language models (LLMs) simulated examinees offer a promising route to early calibration, but their responses are too accurate and too uniform. We propose Cognitive Diagnostic Profiling (CDP), a zero-shot framework that prompts LLMs to simulate plausible examinees with diverse cognitive profiles: binary attribute-mastery patterns are rendered as natural-language profiles and sampled under an uninformative or an informative distribution. Using the Tatsuoka fraction-subtraction dataset (536 examinees, 15 items, five attributes), we evaluated eight LLM configurations under no-profile, uninformative-CDP, and informative-CDP conditions, assessing alignment with human examinees at the ability-distribution, mastery-profile, and item-difficulty levels. CDP improved all three levels: distributional overlap rose across configurations; weighted correlations between profile-level scores and human profile expectations reached 0.92 to 0.98; and item-difficulty recovery improved in rank order and absolute alignment, most for reasoning-enabled models; in the strongest case, Gemini 3.0 Flash (Thinking), one-parameter logistic (1PL) difficulty Spearman correlations rose from 0.24 to 0.86 and 0.90 and the root-mean-square error (RMSE) fell from 6.31 to 1.30 and 0.90; the informative condition helped most where profile-level alignment was strong. CDP brings LLM-simulated examinees into closer psychometric alignment with human examinees, making them practical for operational test development.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑