arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

线性偏差探针何时以及为何失效?大语言模型表示中偏差可检测性的几何与统计理论

When and Why Do Linear Bias Probes Fail? A Geometric and Statistical Theory of Bias Detectability in Large Language Model Representations

Mo Hai, Haifeng Li

arXiv 2609.22337首次发表:更新:

AI 中文总结

针对线性探针在部分人口统计信息下失效的问题,提出几何与统计理论,证明纯度定律、曲率上限和可检测阈值,将偏差审计转化为功效分析以确定样本预算。

AI 中文摘要

线性探针是检测大语言模型隐藏表示中社会偏差的标准工具。然而,所报告的探针准确率几乎完全来自反事实评估,其中每个输入都带有明确的人口统计标记。一旦只有一部分α的输入携带人口统计信息,性能就会急剧下降,而一个弱的探针可能反映的是无偏模型或检测能力不足的检测器。我们发展了一种理论来解决这种模糊性。将表示建模为在曲率为κ的流形上具有马氏距离s的两个类别条件簇,我们证明了:(i) 一个由流形外在半径控制的有限样本泛化界,并匹配√(d/n)的极小极大下界;(ii) 最大线性探针AUC的精确纯度定律,随α严格递增;(iii) 曲率上限:空间形式上的环境弦距离不能超过2/√κ;(iv) 一个可检测性阈值,低于该阈值任何审计都无法将探针输出与随机区分。每个定理都在具有已知真实标签的合成流形以及六个开放权重模型×四个偏差维度上得到验证,其中纯度定律仅使用在α=1处测量的单一交叉拟合估计值ŝ即可预测完整的AUC-α曲线,无需对这些曲线拟合任何参数。该框架将偏差审计转变为功效分析:给定目标纯度和效应量,它规定了进行结论性审计所需的样本预算n(α)。

英文摘要

Linear probing is the standard instrument for detecting social biases in the hidden representations of large language models. Yet reported probe accuracies come almost exclusively from \emph{counterfactual} evaluations in which every input carries an explicit demographic marker. Once only a fraction $α$ of inputs carries demographic information, performance degrades sharply, and a weak probe may reflect either an unbiased model or an underpowered detector. We develop a theory that resolves this ambiguity. Modeling representations as two class-conditional clusters with Mahalanobis separation $s$ on a manifold of curvature $\kap$, we prove: (i) a finite-sample generalization bound governed by the manifold's extrinsic radius with a matching $\smash{\sqrt{\dB/n}}$ minimax lower bound; (ii) an exact purity law for the maximum linear-probe AUC, strictly increasing in $α$; (iii) a curvature ceiling: ambient chordal separation on a space form cannot exceed $2/\sqrt{\kap}$; and (iv) a detectability threshold below which no audit can distinguish probe output from chance. Every theorem is validated on synthetic manifolds with known ground truth and on six open-weight models $\times$ four bias dimensions, where the purity law predicts entire AUC--$α$ curves from a single cross-fitted $\hat s$ measured at $α=1$, with no parameters fitted to those curves. The framework turns bias auditing into a power analysis: given a target purity and effect size, it prescribes the sample budget $n(α)$ for a conclusive audit.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑