发表机构
University of Amsterdam(阿姆斯特丹大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究在ECtHR语料库上比较LegalBERT基线与其公平性正则化变体,发现干预不改变公平性但持续降低解释充分性,证明公平性与解释忠实性可解离,需直接评估公平性。
AI 中文摘要
LegalBERT等Transformer模型在法律决策支持中的应用日益广泛,这引发了对公平性及模型解释透明性的担忧。这些属性通常被分别评估,因此尚不清楚改变公平性的去偏干预是否也会改变解释忠实反映模型推理的程度。本研究在LexGLUE的ECtHR指控违反公约语料库上考察了这一关系。研究将LegalBERT基线模型与一种公平性正则化变体进行比较,后者在微调过程中对刻板印象中的温暖和能力表征进行惩罚。评估涵盖了预测性能、人口统计公平性以及SHAP解释的忠实性,共使用五个随机种子。在性能最优的正则化强度下,该干预并未减少人口统计差异。这一零结果在两种公平性定义及性别轴上的常规词对控制条件下均成立。分类性能基本保持不变。然而,该干预在所有五个种子和三个阈值下持续降低了解释的充分性。一个打乱词对的控制实验重现了这种降低,同时性能和公平性保持不变,表明该效应源于对比性表征正则化,而非特定于温暖和能力结构。结果表明公平性与解释忠实性之间存在解离:解释行为的变化并不必然指示公平性的变化,因此必须直接评估公平性。
英文摘要
Transformer models such as LegalBERT are increasingly used in legal decision support, raising concerns about both fairness and the transparency of model explanations. These properties are usually evaluated separately, leaving open whether a debiasing intervention that changes fairness also changes how faithfully explanations reflect model reasoning. This study investigates that relationship on the ECtHR alleged-violations corpus from LexGLUE. It compares a LegalBERT baseline with a fairness-regularized variant that penalizes stereotypical warmth and competence representations during fine-tuning. The evaluation covers predictive performance, demographic fairness, and SHAP explanation faithfulness across five random seeds. At the performance-optimal regularization strength, the intervention does not reduce demographic disparity. This null result holds across two fairness definitions and a conventional word-pair control on the gender axis. Classification performance is largely unchanged. However, the intervention consistently degrades explanation sufficiency across all five seeds and three thresholds. A shuffled-pair control reproduces this degradation while leaving performance and fairness unchanged, indicating that the effect arises from contrastive representational regularization rather than specifically from the warmth and competence structure. The results demonstrate a dissociation between fairness and explanation faithfulness: changes in explanation behavior do not necessarily indicate changes in fairness, and fairness must therefore be evaluated directly.