AI 中文总结
本研究通过扩展LIAR基准数据集,评估六个LLM的假新闻检测性能,发现存在性别偏见,影响检测的可靠性与公平性,且数据集已公开。
AI 中文摘要
大语言模型(LLMs)越来越多地被用于自动事实核查,然而它们在该场景下的性别偏见问题仍未得到充分探索。本研究首次利用真实世界数据,对基于LLM的假新闻检测中的性别偏见展开系统调查。我们为LIAR基准数据集中的每条陈述,补充了说话人职位名称的三种性别变体(中性、男性、女性),以测试真实性判断是否仅因性别呈现而变化。我们在多个偏见与公平性指标上评估了六个最先进的LLM,所有模型均表现出性别敏感性:9.79%-35.13%的陈述在三种变体间获得不一致标签,其中男女变体间的对比显示出6.5%-23.6%的标签翻转率。我们识别出两种主要的偏见表现形式:不稳定性(判断不一致)与方向性(系统性偏袒)。五个模型呈现出统计学显著的方向性效应,最强效应表现为男性怀疑模式。这些发现表明,性别偏见破坏了基于LLM的假新闻检测的可靠性与公平性,凸显了对偏见感知的评估与缓解策略的需求。本研究补充后的数据集已公开发布,以支持未来研究。
英文摘要
Large Language Models (LLMs) are increasingly used for automated fact-checking, yet their susceptibility to gender bias in this context remains underexplored. This study presents the first systematic investigation of gender bias in LLM-based fake news detection using real-world data. We augment the LIAR benchmark with three gender variants of speaker job titles (Neutral, Male, Female) for each statement to test whether veracity judgments vary solely based on gender presentation. Six state-of-the-art LLMs are evaluated across multiple bias and fairness metrics. All models exhibit gender sensitivity: 9.79%-35.13% of statements receive inconsistent labels across the three variants, with Male-Female comparisons showing 6.5%-23.6% flip rates. Two primary bias manifestations are identified: instability (inconsistent judgments) and directionality (systematic favoritism). Five models show statistically significant directional effects, with the strongest effects displaying male-skeptic patterns. These findings demonstrate that gender bias undermines both reliability and fairness in LLM-based fake news detection, highlighting the need for bias-aware evaluation and mitigation strategies. The augmented dataset is publicly released to support future research.
CommentsAccepted to the 6th Workshop on Bias and Fairness in AI at ECML PKDD 2026. Dataset available at https://github.com/raziehch/GenderedLIARDataset