发表机构
Sigma Jahan(Sigma Jahan)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究Transformer语言模型中偏差输出难以定位和修复的问题,提出白盒头级公平性调试方法ROBIN,通过对注意力头排序并去除偏差子空间来修复偏差,在试点研究中减少WinoBias差距且更好保留语言建模质量。
AI 中文摘要
Transformer语言模型越来越多地用作软件组件,但模型内部的偏差输出仍难以定位和修复。现有公平性测试和修复方法大多在输入-输出或重新训练层面操作,而近期研究表明偏差相关行为可能集中在少数注意力头中。本文研究能否通过有针对性的推理时干预来定位和修复注意力头。我们引入了ROBIN,一种白盒头级公平性调试方法,该方法根据对公平性探针的敏感度对注意力头进行排序,并从选定头的输出中去除小的偏差子空间。在一项四模型试点研究中,ROBIN减少了所有模型中测得的WinoBias差距,同时比全头归零更好地保留了语言建模质量。这些初步结果表明,头级偏差修复不仅应考虑选择哪些头,还应考虑如何修改选定的头。
英文摘要
Transformer language models are increasingly used as software components, yet biased outputs remain difficult to localize and repair inside the model. Existing fairness testing and repair methods largely operate at the input-output or retraining level, while recent work suggests that bias-related behavior can concentrate in a small set of attention heads. This paper studies whether attention heads can be localized and repaired through a targeted inference-time intervention. We introduce ROBIN, a white-box head-level fairness debugging method that ranks attention heads using sensitivity to fairness probes and removes a small bias subspace from selected head outputs. In a four-model pilot study, ROBIN reduces the measured WinoBias gap across all models while preserving language-modeling quality better than whole-head zeroing. These preliminary results suggest that head-level bias repair should consider not only which heads are selected, but also how selected heads are modified.
CommentsAccepted in ICSME NIER track, 2026