发表机构
NASK - National Research Institute; Warsaw University of Technology; Gdańsk University of Technology(NASK国家研究院; 华沙理工大学; 格但斯克理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出基于参考的表示偏差偏移ΔB方法,通过相对表示审计LLM隐藏状态偏差,无需任务数据,与输出级基准相关且计算高效,可检测微调后偏差增加。
AI 中文摘要
现有的偏差审计方法通常依赖于模型输出,需要昂贵的基准或评判模型,并且可能遗漏从未出现在生成文本中的内部变化。我们提出了一种基于参考的方法,用于审计相关模型变体(例如微调前后)隐藏状态表示中的偏差。由于微调重塑了表示几何结构,绝对隐藏状态无法直接比较,因此我们通过每个句子与一组固定锚定句的相似性对其进行编码,在共享的比较空间中生成相对表示。在此空间中,我们衡量目标群体在正面和负面属性关联上的变化,这一量我们称为表示偏差偏移ΔB。在三个模型家族以及WildGuardMix、DecodingTrust和ToxiGen基准上,ΔB与我们测试的18个设置中的15个输出级偏差变化相关,在完全微调下达到|r|=0.84(p<0.001),在参数高效适应下变得更具模型依赖性。对ΔB进行阈值化可检测偏差增加的检查点,ROC AUC在0.65到0.99之间,并且在WildGuardMix和DecodingTrust上,对于所有三个家族,其区分效果优于基于SEAT的基线。ΔB在锚定集、属性集和目标模板变化下也保持稳定。我们的方法不需要任务特定的评估数据,并且在大约三分钟内审计一个模型,使用的计算量比此处考虑的输出级基准少3到50倍。我们认为它是对基于输出的审计的补充,而非替代。
英文摘要
Existing bias auditing methods typically rely on model outputs, requiring costly benchmarks or judge models and potentially missing internal shifts that never appear in generated text. We propose a reference-based method that audits bias in hidden-state representations across related model variants, for example before and after fine-tuning. Because fine-tuning reshapes representation geometry, absolute hidden states are not directly comparable, so we encode each sentence by its similarities to a fixed set of anchor sentences, yielding relative representations in a shared comparison space. There we measure how target groups shift in their association with positive and negative attributes, a quantity we call the Representational Bias Shift $ΔB$. Across three model families and the WildGuardMix, DecodingTrust and ToxiGen benchmarks, $ΔB$ correlates with output-level bias change in 15 of the 18 settings we test, reaching $|r| = 0.84$ ($p < 0.001$) under full fine-tuning and becoming more model-dependent under parameter-efficient adaptation. Thresholding $ΔB$ detects checkpoints whose bias increased with ROC AUC between $0.65$ and $0.99$, and on WildGuardMix and DecodingTrust it separates them better than a SEAT-based baseline for all three families. $ΔB$ is also stable under changes to the anchor set, attribute sets and target templates. Our method requires no task-specific evaluation data and audits a model in about three minutes, using $3$-$50\times$ less compute than the output-level benchmarks considered here. We view it as complementary to output-based auditing rather than a replacement for it.