LoRA微调模型用于控制系统课程问答:模型规模与秩效应的多维评估
LoRA Fine-Tuned Models for Control Systems Course Q\&A: A Multidimensional Evaluation of Model Scale and Rank Effects
- State Key Laboratory of Synthetical Automation for Process Industries(流程工业综合自动化国家重点实验室)
- Northeastern University(东北大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究使用LoRA微调Qwen2.5-3B和7B模型,在控制系统课程问答数据集上评估不同秩的效果,发现LoRA能提升参考答案相似性和结构化输出稳定性,其中7B-r16表现最佳。
AI中文摘要:
大型语言模型(LLMs)越来越多地被用于专业大学课程中,但控制系统问题需要协调的术语、符号、推导和逐步解释。通用的直接回答可能结构不一致且难以验证。我们利用线性控制系统课程中的习题和参考答案,构建了一个包含360个系统-用户-助手对话的监督微调数据集。我们将LoRA应用于Qwen2.5-3B-Instruct和Qwen2.5-7B-Instruct。在相同的数据划分、推理设置和评估协议下,我们比较了基础模型和微调模型,并测试了LoRA秩r=4、8和16。评估使用了ROUGE、BERTScore和结构化输出特征来衡量参考答案相似性和“解法-方法-教学要点”格式的稳定性。LoRA在两种模型规模下均提高了相似性和结构化输出稳定性。在当前测试集上,7B-r16取得了最高的ROUGE-L(0.4093)和BERTScore-F1(0.8643),而r=8在性能和参数效率之间提供了更好的平衡。Bootstrap重采样显示,3B-r16的ROUGE-L提升为0.0764 [0.0613, 0.0915],7B-r16的提升为0.0874 [0.0687, 0.1042];两个区间均超过零,表明在当前测试集上文本相似性有稳定提升。这些结果表明,LoRA可以使开源指令微调模型更紧密地对齐课程参考答案的语言和教学组织方式。然而,这些指标主要捕捉文本相似性和格式一致性,而非领域特定的推理或数学正确性,后者需要专家评估和任务特定的评分标准。
英文摘要:
Large language models (LLMs) are increasingly used in specialized university courses, but control-systems questions require coordinated terminology, notation, derivations, and stepwise explanations. Direct general-purpose responses may be inconsistently structured and hard to verify. Using exercises and reference solutions from a Linear Control Systems course, we built a supervised fine-tuning dataset of 360 system-user-assistant conversations. We applied LoRA to Qwen2.5-3B-Instruct and Qwen2.5-7B-Instruct. With identical data splits, inference settings, and evaluation protocols, we compared base and fine-tuned models and tested LoRA ranks r=4, 8, and 16. Evaluation used ROUGE, BERTScore, and structured-output features to measure reference-answer similarity and stability of the Solution-Method-Teaching Points format. LoRA improved both similarity and structured-output stability at both sizes. On the current test set, 7B-r16 achieved the highest ROUGE-L (0.4093) and BERTScore-F1 (0.8643), while r=8 offered a better balance between performance and parameter efficiency. Bootstrap resampling showed ROUGE-L gains of 0.0764 [0.0613, 0.0915] for 3B-r16 and 0.0874 [0.0687, 0.1042] for 7B-r16; both intervals exceeded zero, indicating stable textual-similarity improvements on the current test set. These results suggest LoRA can align open-source instruction-tuned models more closely with the language and pedagogical organization of course reference answers. However, the metrics mainly capture textual similarity and formatting consistency, not domain-specific reasoning or mathematical correctness, which require expert assessment and task-specific rubrics.