针对学生数学隐喻回答的密码本引导编码任务微调大型语言模型
Fine-Tuning Large Language Models for Codebook-Guided Coding of Students' Mathematics Metaphor Responses
- University of Michigan–Ann Arbor(密歇根大学安娜堡分校)
- University of Delaware(特拉华大学)
- Texas A&M University(德克萨斯农工大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究通过LoRA微调开放权重LLM,使其在学生数学隐喻的密码本引导编码任务中性能提升,表现媲美甚至超越仅提示的专有模型,可用于数学教育的规模化AI辅助测量。
AI中文摘要:
学生生成的数学隐喻能够揭示其态度、信念、身份与经历,但人工专家对这些主题和语义复杂的开放式回答进行编码耗时且难以规模化。本研究探究基于LoRA的大型语言模型(LLM)监督微调能否提升其在学生数学隐喻的密码本引导编码任务中的表现。我们使用包含2265条6-8年级学生对食物和动物隐喻提示的回答的人工编码语料库,指导LLM完成两项编码任务:效价-强度编码以捕捉学生对数学的情感取向方向与强度,以及主题编码以捕捉学生通过隐喻表达的对数学的建构。我们将两个专有模型GPT-4o mini和GPT-5 mini(仅提示条件下)与两个开放权重模型DeepSeek-R1 1.5B和Mistral 7B在微调前后进行对比。结果显示,微调后开放权重模型在两项任务上的性能和运行间可靠性较基础版本均有显著提升,微调后的紧凑型开放权重模型表现可与仅提示的专有模型媲美,且在多数情况下超越后者。这些发现表明,紧凑型开放权重LLM可支持数学教育中对学生隐喻回答的规模化、本地可控且注重隐私的AI辅助测量。
英文摘要:
Student-generated metaphors about mathematics can provide insights into students' attitudes, beliefs, identities, and experiences, but expert human assessment through thematic coding of these semantically complex metaphor responses is labor-intensive and difficult to scale. This study examines whether Low-Rank Adaptation (LoRA)-based supervised fine-tuning of Large Language Models (LLMs) can improve their performance on codebook-guided coding tasks for student mathematics metaphors. We utilized a human-coded corpus of 2,265 Grade 6-8 responses to food- and animal-based metaphor prompts and evaluated LLMs on two tasks: valence-intensity coding of students' affective orientations toward mathematics and thematic coding of their metaphorical framings of mathematics. Two open-weight LLMs, DeepSeek-R1 1.5B and Mistral 7B, were evaluated before and after fine-tuning and compared with two proprietary LLMs, GPT-4o mini and GPT-5 mini. Results show that fine-tuning substantially improved the performance and run-to-run reliability of the open-weight LLMs across both tasks relative to their base versions, making the fine-tuned LLMs competitive with and often outperforming the proprietary LLMs. These findings suggest the potential of fine-tuned open-weight LLMs for scalable and automated AI-assisted measurement of students' metaphor responses with competitive performance while maintaining local controllability and privacy-conscious deployment.