arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37887cs.LG

量化语言模型的行为能力证书

Behavioral Capacity Certificates for Quantized Language Models

Arian Eamaz, Mojtaba Soltanalian

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出行为能力证书(BCC),通过聚合实现质量对量化模型行为收费,支持三步部署流程,并在多个模型上验证了其有效性。

中文摘要 AI 辅助

激活和键值缓存精度会改变量化语言模型的计算结果,而不会改变其存储的权重。然而,直接的权重码界为行为不同的部署分配了相同的复杂度,并为行为相同的权重码分别收费。行为能力证书(BCC)使用完整实现(权重、缩放、激活和缓存规则)的聚合先验质量来对行为收费,这些实现会引发相同的有界损失。当量化合并实现时,这种共享质量会降低复杂度惩罚,并且一个盈亏平衡法则决定了节省是否能在验证成本中幸存。BCC支持三步部署工作流程,我们的实验验证了每一步。首先,前向筛选根据候选扰动保留参考预测的频率来筛选每层位宽,其质量与Hessian引导选择相当,但预处理成本更低。其次,边际认证单元识别可以在不改变部署行为的情况下被剪枝或符号翻转的权重:每个允许的组合都保留所有声明的预测,在OLMoE-1B-7B和SmolLM2-1.7B上,独立探针限制了任何允许组合改变新文本预测的概率。第三,BCC限制了部署模型的总体损失,对完整解码器非空洞地,且比压缩码路径更严格。在相同缓存内存下,给键比值更高的精度在GPT-2、Qwen2.5和SmolLM2上产生更低的负对数似然和更高的预测一致性,同时在GPT-2审计中产生更严格的复杂度界。

英文摘要

Activation and key-value cache precision change what a quantized language model computes without altering its stored weights. Direct weight-code bounds, however, assign identical complexity to deployments that behave differently and charge separately for weight codes that behave identically. Behavioral Capacity Certificates (BCC) charge for behavior using the aggregate prior mass of complete implementations---weights, scales, activation and cache rules---that induce the same bounded loss. When quantization merges implementations, this shared mass lowers the complexity penalty, and a break-even law determines when the saving survives the cost of validating it. BCC supports a three-step deployment workflow, and our experiments verify each step. First, a forward-only screen shortlists per-layer bit-widths by how often candidate perturbations preserve the reference predictions, with quality comparable to Hessian-guided selection at lower preprocessing cost. Second, margin-certified cells identify weights that can be pruned or sign-flipped without changing the deployed behavior: every permitted combination preserves all declared predictions, and on OLMoE-1B-7B and SmolLM2-1.7B, independent probes bound the probability that any permitted combination changes a prediction on new text. Third, BCC bounds the population loss of the deployed model, nonvacuously for complete decoders and more tightly than the compressed-code route. At equal cache memory, giving keys higher precision than values yields lower NLL and higher prediction agreement on GPT-2, Qwen2.5, and SmolLM2, together with a tighter complexity bound in the GPT-2 audit.

发表机构

  • University of Illinois Chicago(伊利诺伊大学芝加哥分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑