发表机构
Tohoku University; National Institute of Informatics; University of Tokyo; Osaka Kyoiku University(东北大学; 国立情报学研究所; 东京大学; 大阪教育大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究利用真实学习者评价数据,检验LLMs在教育反馈场景中模拟学习者评价的能力,发现其能力有限,提供个性化信息可改善个体模拟但难提升群体一致性。
AI 中文摘要
尽管近期研究已探索使用大型语言模型(LLMs)进行人类行为和偏好模拟,但在教育环境中,LLMs 能否模拟真实学习者的主观评价仍不清楚。我们利用真实学习者对高中生物问题反馈的评价数据,在群体和个体两个层面研究这一问题。我们比较了六种模型在有无学习者特定信息(如个性特征和评价示例)时的表现。结果表明,LLMs 模拟学习者评价的能力仍然有限。提供学习者画像和示例可改善评分校准和个体层面的模拟,但往往无法提升群体层面的一致性。这些发现强调了研究哪些学习者信息和适应策略对学习者偏好模拟有效的必要性。
英文摘要
While recent studies have explored human behavior and preference simulation using large language models (LLMs), it remains unclear how well LLMs can simulate subjective evaluations from real learners in educational settings. We investigate this question using real learner evaluation data on feedback for high-school biology questions at both the group and individual levels. We compare performance with and without learner-specific information, such as personality traits and evaluation examples, across six models. Our results show that LLMs still have a limited ability to simulate learner evaluations. Providing learner profiles and examples improves score calibration and individual-level simulation, but more often fails to improve group-level consistency. These findings highlight the need to investigate which learner information and adaptation strategies are effective for learner preference simulation.
CommentsAccepted to the EMNLP 2026 Main Conference