针对各向异性校准WEAT:将ZCA白化作为嵌入关联测试的几何预处理步骤
Calibrating WEAT Against Anisotropy: ZCA Whitening as a Geometric Pre-Processing Step for Embedding Association Tests
浏览论文内容
中文总结 AI 辅助
本研究提出将ZCA白化作为预处理步骤校准WEAT,可降低嵌入空间各向异性,超30%WEAT结果显著性改变,提升语义相似度,为相关偏差测量提供更可靠基础。
中文摘要 AI 辅助
我们提出零相位分量分析(ZCA)白化作为词嵌入关联测试(WEAT)的几何预处理步骤。WEAT是一种广泛应用于计算社会科学和AI公平性研究的偏差测量方法,它依赖余弦相似度作为语义关联的度量,该方法假设嵌入空间近似各向同性。但已有研究表明,许多广泛使用的语言模型不满足这一假设,引发了对偏差测量可靠性的担忧。ZCA白化将嵌入空间的协方差转换为单位矩阵,同时最小化对原始向量的扰动,该转换恢复了WEAT所依赖的各向同性条件。我们在10个标准WEAT测试套件和7个涵盖三个架构系列的模型上评估了所提方法,共得到70个模型-任务组合。结果显示,ZCA白化大幅降低了所有模型嵌入空间的各向异性,对于高度各向异性的模型,我们还观察到标准语义相似度基准的提升,表明校准后的空间能更好地捕捉语义关联。校准后,超过30%的WEAT结果改变了显著性状态,效应量根据偏差类别向两个方向变化。这些变化表明,未校准的测量可能同时高估和低估嵌入空间中编码的关联。研究结果表明,应对各向异性嵌入空间中先前报告的偏差测量谨慎解读,可能需要使用校准方法重新评估,我们的方法有助于在计算社会科学和AI公平性研究中恢复WEAT的测量基础。
英文摘要
We propose Zero-phase Component Analysis (ZCA) whitening as a geometric pre-processing step for the Word Embedding Association Test (WEAT). WEAT is a bias measurement method widely used in both computational social science and AI fairness research. It relies on cosine similarity as a measure of semantic association, which assumes that the embedding space is approximately isotropic. However, prior work has reported that many widely used language models do not satisfy this assumption, raising concerns about the reliability of bias measurements. ZCA whitening transforms the covariance of the embedding space into the identity matrix while minimizing perturbation to the original vectors. This transformation restores the isotropy condition on which WEAT relies. We evaluate our approach on ten standard WEAT test suites and seven models spanning three architectural families, yielding 70 model-task combinations. The results show that ZCA whitening substantially reduces the anisotropy of the embedding spaces across all models. Particularly for highly anisotropic models, we further observe improvements on standard semantic similarity benchmarks, indicating that the calibrated space better captures semantic associations. After calibration, over 30% of WEAT results change significance status, and effect sizes shift in both directions depending on bias category. These shifts suggest that uncalibrated measurements may both overestimate and underestimate the associations encoded in the embedding space. These findings indicate that previously reported bias measurements in anisotropic embedding spaces should be interpreted with caution and may benefit from re-evaluation with calibrated methods. Our approach contributes to restoring the measurement foundation of WEAT across both computational social science and AI fairness research.