适用于CPU可部署安全分类的可复现、许可感知蒸馏方案
A Reproducible, License-Aware Distillation Recipe for CPUDeployable Safety Classification
AI总结:
本文提出一种可复现、许可感知的知识蒸馏方案,训练小型学生模型实现CPU可部署的安全分类,其性能接近大参数教师模型,误报率更低且推理速度更快。
AI中文摘要:
在商用硬件上部署大语言模型的安全层受限于可用的防护措施:当前开源防护模型的参数规模在10亿至90亿之间,面向图形处理器(GPU)设计,在中央处理器(CPU)上处理每个请求耗时达数秒。本文提出一种可复现、许可感知的知识蒸馏方案,以应对该限制。一个强大的开源防护模型对来自24个公开数据集、约97000条提示词的语料进行标注,标注为与公开危害分类体系对齐的7个安全类别;随后训练了一组小型学生模型,涵盖词汇型、浅层、编码器型和生成型架构,以复现该标注信号。语料按许可边界划分,使得可部署模型与研究模型仅在训练数据上存在差异,该限制的成本也因此可量化。所有模型均在包含6361条数据的独立黄金基准上进行评分,该基准分为4个切片,标注过程独立于教师模型,且包含无害提示词切片,可用于衡量过度防护情况。蒸馏后的学生模型在重叠置信区间内与教师模型对抗文本的表现相当,且降低了无害提示词的误报率;最小的生成型学生模型误报率为3.8%,而80亿参数教师模型的误报率为4.8%,编码器模型在CPU上处理每个请求耗时约24毫秒。按类别重新平衡是该方案唯一的关键要素,本文未声称蒸馏后的防护模型具有优越性,在干净参考切片上,它们仍领先于教师模型。
英文摘要:
Deploying a safety layer for large language models on commodity hardware is constrained by the guards available to do it: current open guard models hold between 1 and 9 billion parameters, are oriented toward the graphics processing unit, and answer in seconds per request on a central processing unit. This paper presents a reproducible, license-aware knowledge-distillation recipe addressing that constraint. A strong open guard labels a corpus of roughly 97,000 prompts, drawn from 24 public datasets, into seven safety categories aligned to a public hazard taxonomy, and a fleet of small students spanning lexical, shallow, encoder and generative architectures is trained to reproduce that signal. The corpus is partitioned at the license boundary, so that a deployable and a research model differ only in their training data and the cost of that restriction becomes measurable. Every model is scored against an independent gold benchmark of 6,361 rows over four slices, labeled apart from the teacher and including a slice of harmless prompts that makes over-defense measurable. The distilled students match the teachers on adversarial text within overlapping confidence intervals and reduce false alarms on harmless prompts, the smallest generative student reaching 3.8% against 4.8% for the 8-billion-parameter teacher, while the encoder classifies in roughly 24 ms per request on CPU. Per-class rebalancing is the only decisive ingredient of the recipe. No superiority over the distilled guards is claimed; on the clean reference slice they remain ahead.