发表机构
University of Galway; School of Computer Science(戈尔韦大学; 计算机科学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究小型语言模型多步算术推理困难问题,通过生成结构化合成推理数据并结合LoRA在消费硬件上微调Qwen3模型,提高了模型在GSM8K等基准测试上的准确率和迁移效果,证明低成本合成数据设计可改善小型语言模型算术适应性。
AI 中文摘要
小型语言模型在本地部署中很有吸引力,但在多步算术推理方面存在困难。我们研究在消费硬件约束下,结构化合成推理数据能否改善这种情况。从GSM8K开始,我们使用GPT - 5 - mini生成了一个包含21,250个小学数学文字问题变体的语料库,然后在消费硬件上对Qwen3 - 0.6B和Qwen3 - 1.7B进行LoRA微调。结果显示,模型在GSM8K上的精确匹配准确率提高,在相关算术基准测试上的迁移效果更好。定性分析表明微调模型有诸多优势。这些结果表明低成本合成数据设计可显著改善小型语言模型的算术适应性。
英文摘要
Small language models are attractive for local deployment, but they often struggle with multi-step arithmetic reasoning. We study whether structured synthetic reasoning data can improve this behaviour under consumer-hardware constraints. Starting from GSM8K, we generated a 21,250-example corpus of grade-school arithmetic word-problem variants using GPT-5-mini, combining natural-language solution traces, light Socratic-style cues, structural variation, and irrelevant distractor context. We then fine-tuned Qwen3-0.6B and Qwen3-1.7B with LoRA on consumer hardware (Apple M4, 16 GB RAM). Exact-match accuracy on GSM8K improved from 36.5% to 49.1% for Qwen3-0.6B and from 53.5% to 66.5% for Qwen3-1.7B. For Qwen3-1.7B, transfer to related arithmetic benchmarks was stronger, reaching 98.9% on MultiArith and 73.0% on SVAMP, compared with 54.4% and 45.3% for the base model. Qualitative analysis suggests that fine-tuned models produce shorter reasoning traces, make fewer arithmetic and distractor-use errors, and benefit more consistently from self-consistency sampling. These results show that low-cost synthetic data design can materially improve arithmetic adaptation in small language models. Because the intervention combines Socratic-style cues with other data-design choices, we interpret the gains as evidence for structured synthetic reasoning data rather than as a causal test of Socratic guidance alone.