AI 中文总结
本研究审计ISOT假新闻语料库,发现高准确率源于主题和来源泄漏而非真实性,线性模型在主题偏移下更稳健,建议使用廉价基线诊断捷径学习。
AI 中文摘要
在ISOT/Kaggle“假新闻与真实新闻”语料库上训练出的文本分类器通常报告准确率和F1分数高于0.98,这种性能水平与评估真实性的难度相比显得格格不入。我们使用一个透明的TF-IDF和线性分类器流程作为测量工具,沿着三条泄漏通道和两种分布偏移协议对该语料库进行审计,并发布了所有代码和派生数据。首先,该基准部分退化:仅给定主题元数据字段、丢弃文章文本的分类器达到了F1=1.000,因为两个类别具有不相交的主题。其次,移除所有三条泄漏通道——元数据、存在于99.2%真实文章中的新闻通讯社来源标签,以及污染了朴素测试集19.4%的6,251篇重复文档——仅使F1下降1.21个百分点(从0.9935降至0.9814);残余信号是分散的编辑风格而非少数泄露标记,因为删除权重最高的1,000个一元词组后F1仍为0.926。第三,这种风格信号不可迁移:在主题不相交协议下,平均精度从0.9995降至0.9475,部署F1从0.9905降至0.8067,先验匹配分析确认了真实的5.2个百分点的判别力损失,而时间迁移几乎无损。微调的DistilBERT在分布内更强(F1=0.9993),但在主题偏移下退化严重得多,平均精度损失12.9个百分点,而线性模型损失5.2个百分点。迁移到独立的LIAR基准上,所有三个模型均降至接近随机排序水平(ROC-AUC 0.54-0.57),无一超过多数类基线。我们得出结论:该语料库内的分数量化的是来源和主题的可分离性而非真实性,增加模型容量利用的是捷径而非规避捷径,并建议将仅元数据、小样本和主题不相交基线作为未来工作的廉价诊断工具。
英文摘要
Text classifiers trained on the ISOT/Kaggle "Fake and Real News" corpus routinely report accuracy and F1 above 0.98, a level of performance that sits uneasily beside the difficulty of assessing veracity. Using a transparent TF-IDF and linear-classifier pipeline as a measurement instrument, we audit the corpus along three leakage channels and two distribution-shift protocols, releasing all code and derived numbers. First, the benchmark is partly degenerate: a classifier given only the subject metadata field, with the article text discarded, attains F1 = 1.000, since the two classes have disjoint subjects. Second, removing all three leakage channels, metadata, a newswire source tag present in 99.2% of real articles, and 6,251 duplicate documents contaminating 19.4% of a naive test split, lowers F1 by only 1.21 points (0.9935 to 0.9814); the residual signal is diffuse editorial style rather than a few giveaway tokens, since deleting the 1,000 highest-weight unigrams still leaves F1 = 0.926. Third, this style signal does not transfer: under a topic-disjoint protocol, average precision falls from 0.9995 to 0.9475 and deployed F1 from 0.9905 to 0.8067, with a prior-matched analysis confirming a genuine 5.2-point loss of discrimination, while temporal transfer is nearly lossless. A fine-tuned DistilBERT is stronger in-distribution (F1 = 0.9993) but degrades far more under topic shift, losing 12.9 average-precision points against the linear model's 5.2. Transferred to the independent LIAR benchmark, all three models fall to near-chance ranking (ROC-AUC 0.54-0.57), none beating a majority-class baseline. We conclude that within-corpus scores here quantify source and topic separability rather than veracity, that added capacity exploits the shortcut rather than avoiding it, and we recommend metadata-only, small-sample, and topic-disjoint baselines as inexpensive diagnostics for future work.
Comments17 pages, 11 figures, 11 tables. Code, experiment scripts, and machine-readable results: https://github.com/vermayuvraj/fake-news-detection