arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

压缩语言模型的验证:审计与统计工具包

Certifying Compressed Language Models: An Audit and a Statistical Toolkit

Amogh Singh

arXiv 2608.15046首次发表:更新:

发表机构

Georgia Institute of Technology(佐治亚理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对压缩语言模型等价声明的证据不足问题,开展17项等价声明的预注册审计,提出配对等价测试的报告标准,发布相关输出与代码。

AI 中文摘要

基准模型精度的微小差值通常被作为压缩模型与原模型等价的证据,但该指标在两个模型最相似时信息含量最低:净差值是单个样本相反变化抵消后剩余的部分,而抵消在等价声明所处的场景中最为完全。本文从公开的单个样本评估数据中挖掘出1707个模型-任务配对单元(规模从13亿到4050亿参数),发现 churn(单个样本的变化量)约为净精度差值的5倍,且与基线得分相同的单元在单个样本层面仍存在差异。本文对来自三类注册框架(方法论文、模型卡片、供应商文档)的17项等价声明进行预注册审计,其中16项符合审计条件;所有声明均未声明预期的数值等价容差,也未发布任务匹配的单个样本输出,仅3项发布了其他任务的输出,5项报告信息过少无法进行数值评估,读者无法在任何样本量下验证这些声明。本文审计的是证据充分性而非真实性,不判定任何声明为假。本文提供了缺失的工具:在声明的容差下进行配对等价测试,生成验证表,该表基于压缩过程中观察到的分歧而非独立二项方差计算评估所需的单个样本。控制实验在字节相同的校准样本上配对GPTQ和AWQ,跨越5个随机种子;在固定的8单元决策规则下,H3假设得到支持:改变校准样本足以在8个验证单元中的5个中反转观察到的方法排序。本文提出的报告标准为5项内容:声明容差、运行配对测试、报告churn与净差值、说明达到的样本量、发布单个样本输出;该标准适用于任何两个相似到值得比较的模型之间的比较。所有单个样本输出、协议和代码均已发布。

英文摘要

A fraction of a point of benchmark accuracy is the usual evidence that a compressed model is equivalent to its original. That quantity is least informative when two models are most alike: a net delta is what survives cancellation between opposing per-item changes, and cancellation is most complete in the regime equivalence claims occupy. Across an atlas of 1,707 paired model-by-task cells mined from public per-item evaluation dumps (1.3B-405B), churn runs roughly five times the net accuracy delta, and cells scoring identically to their baseline still disagree on individual items. In a preregistered audit of 17 equivalence claims from three registered frames (method papers, model cards, vendor documentation), 16 are eligible. None states a prospective numerical equivalence margin, and none releases task-matched per-item outputs, though 3 release outputs for other tasks only; 5 report too little to assess numerically, so a reader cannot check them at any sample size. We audit evidential sufficiency, not truth: no claim is called false. We supply the missing instrument: paired equivalence testing at a declared margin, with certification tables giving the items an evaluation needs, computed from disagreement observed under compression, not from independent-binomial variance. A controlled experiment pairs GPTQ and AWQ on byte-identical calibration samples across five seeds. Under the frozen eight-cell decision rule H3 is supported: changing the calibration draw was sufficient to reverse the observed method ordering in 5 of 8 confirmatory cells. The reporting standard we propose is five lines: declare a margin, run the paired test, report churn beside net delta, cite the sample size you met, release per-item outputs. It applies to any comparison between two models alike enough to be worth comparing. All per-item outputs, protocols and code are released.

Comments109 pages, 3 figures, 20 tables. Artifacts and per-item outputs: doi.org/10.5281/zenodo.21939143

DOI:10.5281/zenodo.21939143

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑