余弦相似度不是证据:量化下可解释性迁移的噪声基底测量
Cosine Similarity Is Not Evidence: Measuring the Noise Floor of Interpretability Transfer Under Quantization
浏览论文内容
中文总结 AI 辅助
本研究指出,量化下可解释性迁移的余弦相似度报告缺乏噪声基底,无法作为证据;通过测量类别分离度并对比分裂半样本零分布,发现INT4方向旋转而INT8无移动,并提出低成本报告建议。
中文摘要 AI 辅助
一个在没有解释所需数量的情况下报告的统计量不是证据。我们针对人工智能安全中的一个具体实践展开这一论点。可解释性产物在全精度权重上校准,在量化权重上部署,并通过尺度不变统计量(余弦相似度、相关性、AUROC)被证明能在这种变化中幸存,而这些统计量在报告时未附带其噪声基底。对于均值差方向估计器,分裂半样本噪声基底由一个无量纲数 κ = nρ²/d 控制。闭式表达式 E[cos] ≈ (1+4/κ)^{-1} 是经典的;缺失的输入是类别分离度 ρ,我们在真实激活上测量了它;据我们所知,没有一项压缩迁移研究报告过该值。在 Qwen2.5-1.5B-Instruct 上,ρ 在深度上为 33 至 61,因此估计器的两次独立运行仅因采样就能达到 0.978 至 0.994 的一致性。因此,已发表的全精度与量化拒答方向之间的余弦值 0.996 不能在没有计算该值所用的 n(未报告)的情况下被解读为保留。在 n 已知的情况下,我们根据在该量化模型内测量的分裂半样本零分布来判断每个低位余弦,因为全精度零分布假设低位估计器具有相同的方差。这个假设正是零分布存在的意义所在。结果很明确:在 INT4 下方向发生了旋转,且偏差超过了估计器自身的噪声。在 INT8 下我们未检测到移动,这并非等价性声明。我们还表明,尺度不变统计量无法区分迁移决策变量的平移与衰减,尽管两者需要相反的处理措施。最后我们提出报告建议,其成本仅为一次前向传播。代码、数据及单单元格复现已在此 https URL 发布。
英文摘要
A statistic reported without the quantity needed to interpret it is not evidence. We develop that thesis for a concrete practice in AI safety. Interpretability artifacts are calibrated on full-precision weights, deployed on quantized ones, and certified as surviving the change by scale-invariant statistics (cosine similarity, correlation, AUROC) that are reported without their noise floor. For the difference-in-means direction estimator, the split-half floor is governed by one dimensionless number, $κ= nρ^2/d$. The closed form $\mathbb{E}[\cos] \approx (1+4/κ)^{-1}$ is classical; the missing input is the class separation $ρ$, which we measure on real activations; no compression-transfer study we know of reports it. On Qwen2.5-1.5B-Instruct, $ρ= 33$--$61$ across depth, so two independent runs of the estimator agree to $0.978$--$0.994$ by sampling alone. A published cosine of $0.996$ between full-precision and quantized refusal directions therefore cannot be read as preservation without the $n$ it was computed at, which is not reported. Where $n$ is known, we judge each low-bit cosine against the split-half null measured within that quantized model, because a full-precision null assumes the low-bit estimator has the same variance. That assumption is exactly what a null exists to test. The result is plain: at INT4 the direction rotated, and the deficit exceeds the estimator's own noise. At INT8 we detect no movement, which is not an equivalence claim. We also show that a scale-invariant statistic cannot distinguish translation from attenuation of a transferred decision variable, although the two call for opposite remedies. We close with reporting recommendations that cost one forward pass. Code, data, and a one-cell reproduction are released at https://github.com/pvarshh/quantinterp
发表机构
- University of Michigan(密歇根大学)
机构由 AI 辅助整理,请以论文原文为准。