发表机构
Binghamton University(宾汉姆顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究在硬件和治理约束下人工智能民主化问题,通过对九个小语言模型在特定基准测试上控制评估,采用共享参数高效微调管道,得出有序工作流程让部分低于3B的模型可成为结构化小众工作负载本地专家的结论。
AI 中文摘要
人工智能民主化并非主要关乎前沿规模的通用性匹配,而是在普通机构实际能满足的硬件和治理约束下,能否选择、审核及专门化有能力的模型。本文通过对九个参数在135M至3B之间的开放权重语言模型进行控制评估来研究该问题,该评估基于为结构化本地部署设计的1085个示例、16个主题的多项选择基准测试,强调符号精度等。然后采用共享的参数高效微调管道,在NVIDIA L4级预算下用4位NF4量化和DoRA/LoRA式适配器对部分模型进行适配。基础评估中,Qwen Coder 3B以75.67%的严格准确率领先,在共享的108个示例的留出微调分割中,适配使多个模型准确率提升。研究得出结论:有序的工作流程使部分低于3B的模型可成为结构化小众工作负载的本地专家。
英文摘要
AI democratization is not primarily a question of matching frontier-scale generality; it is a question of whether capable models can be selected, audited, and specialized under hardware and governance constraints that ordinary institutions can actually satisfy. This paper studies that problem through a controlled evaluation of nine open-weight language models between 135M and 3B parameters on a 1,085-example, 16-topic multiple-choice benchmark designed for structured local deployment. The benchmark emphasizes symbolic precision, constrained formatting, extraction, and short-horizon semantic decision making under a strict one-letter output protocol. A shared parameter-efficient fine-tuning pipeline then adapts a subset of models using 4-bit NF4 quantization with DoRA/LoRA-style adapters on an NVIDIA L4-class budget. In base evaluation, Qwen Coder 3B leads at 75.67% strict accuracy, followed by Qwen2.5 1.5B at 67.10%, Qwen3.5 2B at 64.98%, and Granite 3.3 2B at 64.61%. On the shared 108-example held-out fine-tuning split, adaptation improves Qwen Coder 3B by +26.85 points, SmolLM2 1.7B by +25.92, Qwen2.5 1.5B by +19.44, SmolLM2 360M by +10.18, and SmolLM2 135M by +5.55. Across ranking, topic-level heterogeneity, difficulty strata, failure composition, efficiency frontiers, and topic-conditioned transfer, the same conclusion recurs: a disciplined workflow of benchmark construction, cross-model evaluation, and low-cost specialization already makes a subset of sub-3B models viable as local experts for structured niche workloads.