arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

媒体偏见检测中可解释性的多维评估

A Multi-Dimensional Evaluation of Explainability in Media Bias Detection

Ting Chen, Raina Zhang, Benjamin M. Ampel, Sagar Samtani

arXiv 2607.19954首次发表:更新:

发表机构

Carnegie Mellon University; Indiana University; Georgia State University(卡内基梅隆大学; 印第安纳大学; 佐治亚州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究媒体偏见检测中可解释性,以BABE数据集对BERT和RoBERTa沿预测性能、解释合理性、机制忠实性三个轴评估,还研究注意力监督微调,发现不同架构在这些方面有差异,三者应分别评估。

AI 中文摘要

自动检测媒体偏见很困难,因为偏见框架往往很微妙,而在新闻分析等领域,仅有准确预测是不够的,还需要反映模型潜在推理的解释。我们使用专家偏见注释(BABE)数据集对基于编码器的媒体偏见检测中的可解释性进行了多维评估。具体而言,我们沿着预测性能、解释合理性(与专家理由的词元级对齐)和机制忠实性(在反事实理由屏蔽下紧凑的注意力头集是否恢复预测信号)这三个互补轴研究了BERT和RoBERTa(基础和大型变体)作为分类器。为了引入合理性的变化,我们还研究了注意力监督微调,它将专家理由注释作为辅助训练信号。注意力监督作为对归因合理性的一种干预,而归因方法的有效性在不同架构之间有很大差异。电路分析进一步揭示了不同架构在机制可恢复性方面的显著差异,表明仅模型规模并不能决定电路可压缩性。我们的研究结果表明,预测性能、归因合理性和机制忠实性表征了模型行为的不同方面,在研究媒体偏见检测的可解释性时应分别进行评估。

英文摘要

Detecting media bias automatically is difficult because biased framing is often subtle, yet in domains such as news analysis, accurate predictions alone are insufficient without explanations that reflect the model's underlying reasoning. We present a multi-dimensional evaluation of explainability in encoder-based media bias detection using the Bias Annotations By Experts (BABE) dataset. Specifically, we study BERT and RoBERTa as classifiers (base and large variants) along three complementary axes: predictive performance, explanation plausibility (token-level alignment with expert rationales), and mechanistic faithfulness (whether compact sets of attention heads recover predictive signal under counterfactual rationale masking). To induce variation in plausibility, we additionally investigate attention-supervised finetuning, which incorporates expert rationale annotations as an auxiliary training signal. Attention supervision serves as an intervention on attribution plausibility, while the effectiveness of attribution methods varies substantially across architectures. Circuit analysis further reveals substantial variation in mechanistic recoverability across architectures, suggesting that model scale alone does not determine circuit compressibility. Taken together, our findings suggest that predictive performance, attribution plausibility, and mechanistic faithfulness characterize different aspects of model behavior and should be evaluated separately when studying explainability in media bias detection.

Comments12 pages, 6 figures, under review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑