大语言模型生成的回答能否支持测评开发?一个人工校准的Rasch基准
Can Large Language Model-Generated Responses Support Assessment Development? A Human-Calibrated Rasch Benchmark
浏览论文内容
中文总结 AI 辅助
本研究通过人工校准的Rasch基准评估LLM生成回答在测评开发中的可用性,发现其高内部一致性但低端类别使用不足,无法替代人类试点数据。
中文摘要 AI 辅助
大语言模型(LLMs)被提议作为试点测试的合成受访者,但其有用性取决于它们是否提供测评开发所需的证据。我们在来自6,245名成年人的14个数字使用技能项目上校准了评定量表模型,并使用人类项目参数来评估为1,300个人口统计匹配的人设生成的回答。LLM回答具有高内部一致性($\alpha \approx .94$),但在0.2-0.3%的回答中使用了最低类别,而人类为13.7-21.5%,且没有人设在所有项目上都选择该类别。这些差距改变了两个子量表的回答类别判断和PC目标判断;在人工校准的量表上,14个项目中有12个的infit低于0.70,表明回答比Rasch模型预期的更可预测。预先注册的对提示、类别顺序和人设信息的更改并未恢复人类的下限范围。高内部一致性不足以证明LLM回答可以替代人类试点数据。
英文摘要
Large language models (LLMs) are proposed as synthetic respondents for pilot testing, but their usefulness depends on whether they supply the evidence assessment development requires. We calibrated rating scale models on 14 digital-use skill items from 6,245 adults and used the human item parameters to evaluate responses generated for 1,300 demographically matched personas. LLM responses had high internal consistency ($α\approx .94$) but used the lowest category in 0.2-0.3% of responses versus 13.7-21.5% for humans, and no persona chose it on every item. These gaps changed the response-category judgment on both subscales and the PC targeting judgment; on the human-calibrated scales, 12 of 14 items had infit below 0.70, indicating responses more predictable than the Rasch model expects. Preregistered changes to the prompt, category order, and persona information did not restore the human lower range. High internal consistency is insufficient evidence that LLM responses can replace human pilot data.
发表机构
- Korea University(高丽大学)
机构由 AI 辅助整理,请以论文原文为准。