更准确还是更高效?评估用于数学推理的本地部署紧凑型开放权重语言模型
More Accurate or More Efficient? Evaluating Locally Deployed Compact Open-Weight Language Models for Mathematical Reasoning
查看机构详情
- Florida Institute of Technology(佛罗里达理工学院)
- College of Engineering and Science(工程与科学学院)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本文提出了一种用于评估本地部署语言模型数学推理性能的受控流程,研究三款紧凑型开放权重模型后发现,无单一模型占优,仅靠准确性不足以选择本地模型。
中文摘要 AI 辅助
由于隐私、成本和可访问性方面的原因,大型语言模型正越来越多地部署在本地硬件上。然而,许多评估都侧重于准确性,较少有研究量化本地运行时间和能耗、表征失败模式,或在受控条件下进行配对统计比较。本文提出了一种受控且可追溯的流程,用于评估本地托管的大型语言模型在数学推理任务中的表现,该流程结合了固定推理设置、分层答案提取与验证、明确的失败模式分类,以及按问题划分的资源测量,并通过配对显著性检验和效应量报告准确性。我们在一项初步研究中对三个参数规模均在50亿参数以内的紧凑型开放权重模型进行了验证,分别是Google的Gemma3:4b、Microsoft的Phi3:3.8b和Alibaba的Qwen3:4b,研究涵盖了8年级数学、微积分I以及高级概率与统计三类数据集。所有模型均在同一工作站上通过相同的本地推理服务器运行,使用共享提示模板、受控设置以及每个数据集对应的匹配问题集。没有任何单一模型占据绝对优势:Qwen3:4b在两个数据集上准确性最高,Gemma3:4b在微积分I数据集上表现最佳;但在所有数据集上,Gemma3:4b每瓦时返回的正确答案数约为Qwen3:4b的三倍,且生成的输出令牌数量少得多,而Qwen3:4b每个问题所需的生成时间、能耗和输出量都显著更高。Phi3:3.8b在所有三个数据集上的准确性都显著较低,其较低的提取失败率表明错误答案源于未被解析的输出,不过我们提醒需注意提示格式可能产生的影响。这些初步发现表明,仅以准确性作为选择本地模型的依据是不够的。
英文摘要
Large language models are increasingly deployed on local hardware for privacy, cost, and accessibility reasons. Yet many evaluations emphasize accuracy while fewer quantify local runtime and energy, characterize failure modes, or apply paired statistical comparisons under controlled conditions. This paper presents a controlled, documented procedure for evaluating locally hosted LLMs on mathematical reasoning. It combines fixed inference settings, hierarchical answer extraction and verification, explicit failure-mode classification, and per-question resource measurement, and reports accuracy with paired significance tests and effect sizes. We demonstrate it in a preliminary study of three compact open-weight models under five billion parameters, Gemma3:4b (Google), Phi3:3.8b (Microsoft), and Qwen3:4b (Alibaba), across datasets spanning Grade 8 Math, Calculus I, and Advanced Probability and Statistics. All models ran through the same local inference server on one workstation, using a shared prompt template, controlled settings, and a matched question set per dataset. No single model dominates. Qwen3:4b is most accurate on two datasets and Gemma3:4b on Calculus I, yet Gemma3:4b returns roughly three times more correct answers per watt-hour than Qwen3:4b on every dataset while generating far fewer output tokens; Qwen3:4b requires substantially more generation time, energy, and output per question. Phi3:3.8b is substantially less accurate on all three datasets; its low extraction-failure rate indicates incorrect answers rather than unparsed output, though we caveat possible prompt-format effects. These preliminary findings indicate that accuracy alone is an insufficient basis for selecting a local model.