洞察学生的思维:在LLM学生模拟器中联合建模潜在推理与行动
INSIDE the Student's Mind: Jointly Modeling Latent Reasoning and Action in LLM Student Simulators
浏览论文内容
中文总结 AI 辅助
该研究针对LLM学生模拟器仅复现学生行动却未捕捉其潜在推理的问题,提出INSIDE框架,通过在配对思考轨迹与行动上微调LLM,提升了模拟行动保真度与推理对齐度,最高对齐度达57.9%。
中文摘要 AI 辅助
基于大语言模型(LLM)的模拟器常能复现可观测的行动,却无法捕捉行动背后的潜在推理。在教育领域,学生模拟被越来越多地用于辅导系统评估等各类应用,这一差距尤为显著:两名学生可能因完全不同的原因提交完全相同的内容。我们提出INTERNAL STUDENT DIALOGUE(INSIDE,内部学生对话),这一学生建模框架不仅对LLM进行微调使其表现得像学生,还使其思考得像学生。INSIDE生成基于布鲁姆教育目标分类法的内部对话,涵盖认知、情感和行动维度,并在配对的思考轨迹与行动上微调模型。我们以不同的提示框架作为基线,从两个维度评估:模拟行动的保真度与生成的内部对话的质量。评估显示,INSIDE在行动保真度(匹配真实学生的代码生成)和推理对齐方面均提升了模拟保真度,实现了最高的模型间对齐度,达57.9%。
英文摘要
Large Language Model (LLM)-based simulators often reproduce observable actions but fail to capture the underlying reasoning behind them. In education, where student simulation is increasingly used for various applications such as evaluating tutoring systems, this gap is especially pronounced. Two students may submit identical submissions for entirely different reasons. We present INTERNAL STUDENT DIALOGUE (INSIDE), a student modeling framework that fine-tunes LLMs not only to act like students but also to think like them. INSIDE generates internal dialogue grounded in Bloom's Taxonomy across cognitive, affective, and action dimensions, and fine-tunes models on paired think traces and actions. We baseline against different prompting frameworks and evaluate on two axes: fidelity of simulated actions and quality of generated internal dialogue. Our evaluations show that INSIDE improves simulation fidelity in both action fidelity, matching code generation of real students, and reasoning alignment, achieving the highest alignment across models up to 57.9%.
发表机构
- University of California, Berkeley(加利福尼亚大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。