arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过小语言模型实现人工智能民主化:面向本地部署的结构化基准测试和参数高效微调

Democratizing AI with Small Language Models: Structured Benchmarking and Parameter-Efficient Fine-Tuning for Local Deployment

Daniel Cersosimo

arXiv 2607.16202首次发表:更新:

发表机构

Binghamton University(宾汉姆顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究在硬件和治理约束下人工智能民主化问题,通过对九个小语言模型在特定基准测试上控制评估,采用共享参数高效微调管道,得出有序工作流程让部分低于3B的模型可成为结构化小众工作负载本地专家的结论。

AI 中文摘要

人工智能民主化并非主要关乎前沿规模的通用性匹配,而是在普通机构实际能满足的硬件和治理约束下,能否选择、审核及专门化有能力的模型。本文通过对九个参数在135M至3B之间的开放权重语言模型进行控制评估来研究该问题,该评估基于为结构化本地部署设计的1085个示例、16个主题的多项选择基准测试,强调符号精度等。然后采用共享的参数高效微调管道,在NVIDIA L4级预算下用4位NF4量化和DoRA/LoRA式适配器对部分模型进行适配。基础评估中,Qwen Coder 3B以75.67%的严格准确率领先,在共享的108个示例的留出微调分割中,适配使多个模型准确率提升。研究得出结论:有序的工作流程使部分低于3B的模型可成为结构化小众工作负载的本地专家。

英文摘要

AI democratization is not primarily a question of matching frontier-scale generality; it is a question of whether capable models can be selected, audited, and specialized under hardware and governance constraints that ordinary institutions can actually satisfy. This paper studies that problem through a controlled evaluation of nine open-weight language models between 135M and 3B parameters on a 1,085-example, 16-topic multiple-choice benchmark designed for structured local deployment. The benchmark emphasizes symbolic precision, constrained formatting, extraction, and short-horizon semantic decision making under a strict one-letter output protocol. A shared parameter-efficient fine-tuning pipeline then adapts a subset of models using 4-bit NF4 quantization with DoRA/LoRA-style adapters on an NVIDIA L4-class budget. In base evaluation, Qwen Coder 3B leads at 75.67% strict accuracy, followed by Qwen2.5 1.5B at 67.10%, Qwen3.5 2B at 64.98%, and Granite 3.3 2B at 64.61%. On the shared 108-example held-out fine-tuning split, adaptation improves Qwen Coder 3B by +26.85 points, SmolLM2 1.7B by +25.92, Qwen2.5 1.5B by +19.44, SmolLM2 360M by +10.18, and SmolLM2 135M by +5.55. Across ranking, topic-level heterogeneity, difficulty strata, failure composition, efficiency frontiers, and topic-conditioned transfer, the same conclusion recurs: a disciplined workflow of benchmark construction, cross-model evaluation, and low-cost specialization already makes a subset of sub-3B models viable as local experts for structured niche workloads.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑