AI 中文总结
研究乌尔都语假新闻检测的跨数据集泛化,用XLM-RoBERTa在两个数据集上实验,发现跨数据集转移存在不对称性,Ax-to-Grind数据集中长度混淆影响性能,提出结合双向转移分析和预测崩溃检查的诊断方法来识别混淆驱动行为。
AI 中文摘要
尽管全球有超过2.31亿人说乌尔都语,但乌尔都语假新闻检测仍缺乏资源。此前工作在单个乌尔都语数据集上有很强的域内性能,但跨数据集泛化很少受到系统关注。本文首次对乌尔都语假新闻检测进行跨数据集泛化研究,使用两个公开可用的平衡数据集。在四种实验条件下微调xlm-roberta-base,与使用逻辑回归和支持向量机的TF-IDF基线进行比较。实验发现显著不对称性,从Notri-Fact到Ax-to-Grind的转移宏观F1为0.771,反之则降至0.005。证明这种崩溃源于Ax-to-Grind中系统的长度混淆,假文章平均117词,真文章平均35词。长度消融实验证实混淆会夸大但非唯一驱动域内性能。提供了一种可重复使用的诊断方法,结合双向转移分析和预测崩溃检查来识别多语言假新闻检测设置中由混淆驱动的行为。
英文摘要
Cross-dataset generalisation is a fundamental requirement for deploying text classifiers in real-world settings, yet systematic evaluation across corpora from different sources remains uncommon in fake news detection and virtually absent in sarcasm detection research. This paper presents a unified empirical study of zero-shot cross-dataset transfer in three domains: Urdu fake news detection (FND), English FND, and sarcasm detection. For each domain, we fine-tune xlm-roberta-base on one corpus and evaluate it on a second corpus from a different source, comparing against TF-IDF baselines with Logistic Regression (LR) and Support Vector Machines (SVM). In Urdu FND (Ax-to-Grind vs. Notri-Fact), we identify a severe length confound in the Ax-to-Grind dataset, fake articles average 3.4 times more words than real articles, causing catastrophic A to B transfer collapse (macro F1 = 0.005) while B to A achieves F1 = 0.771. Extension to English FND (WELFake vs. ISOT) and sarcasm detection (TweetEval Irony vs. Sarcasm Corpus V2) reveals that such failure modes extend beyond Urdu, confirming that shortcut learning from distributional artefacts is a cross-lingual, cross-domain challenge in binary text classification. We provide a reusable diagnostic methodology, combining class-conditional length analysis, bidirectional transfer asymmetry, and predicted label collapse inspection, applicable across any binary text classification setting.
Comments18 pages, 18 figures, 7 tables. Extended version of arXiv:2607.14131