arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.18429cs.CRcs.AIcs.CYcs.LG

网络钓鱼邮件检测的对抗鲁棒性:TF-IDF + 逻辑回归与微调DistilBERT的比较研究

Adversarial Robustness of Phishing Email Detection: A Comparative Study of TF-IDF + Logistic Regression and Fine-Tuned DistilBERT

Tanveer Ahmed, Seyedali Pourmoafil

首次发表
浏览论文内容

中文总结 AI 辅助

研究网络钓鱼邮件检测的对抗鲁棒性,对TF-IDF + 逻辑回归和微调DistilBERT两种方法在正常、合成及对抗条件下对比。发现干净数据准确率不能预测对抗鲁棒性,两种模型依赖不同证据但脆弱性相似,对抗测试应成检测评估标准。

中文摘要 AI 辅助

网络钓鱼邮件始终是最顽固的网络安全威胁之一,机器学习分类器被广泛用于检测此类邮件。然而,大多数已报道的检测准确率是在干净的、符合分布的测试数据上测量的,而非针对故意篡改以逃避检测的邮件。本文报告了两种网络钓鱼检测方法的对照、成对比较:一种是TF-IDF + 逻辑回归基线,另一种是在从六个公共数据集抽取的82,255封电子邮件的统一语料库上训练的微调DistilBERT变压器,并在三种条件下进行评估:正常分布内、合成网络钓鱼和对抗性网络钓鱼。两种模型在干净数据上的准确率均超过98%,但在对抗性测试下急剧下降:TF-IDF + LR降至64.00%(下降34.59个百分点),DistilBERT降至63.64%(下降35.40个百分点),差距仅为0.36个百分点。LIME、SHAP和注意力展开分析表明,两种模型依赖不同证据但显示出相似的脆弱性。成对错误分析表明,模型在54.9%的对抗性样本上达成一致,但各自产生了相似数量的排他性错误(分别为24个和25个),表明部分是互补而非相同的失败模式。结果表明,干净数据准确率不能预测对抗鲁棒性,对抗性测试应成为网络钓鱼检测评估的标准组成部分。

英文摘要

Phishing emails remain one of the most persistent cybersecurity threats, and machine-learning classifiers are widely used to detect them. Most reported detection accuracies, however, are measured on clean, in-distribution test data rather than on emails deliberately altered to evade detection. This paper reports a controlled, pairwise comparison of two phishing-detection approaches a TF-IDF + Logistic Regression baseline and a fine-tuned DistilBERT transformer trained on a unified corpus of 82,255 emails drawn from six public datasets and evaluated under three conditions: normal in-distribution, synthetic phishing, and adversarial phishing. Both models exceeded 98% accuracy on clean data yet degraded sharply under adversarial testing: TF-IDF + LR fell to 64.00% (a 34.59-percentage-point drop) and DistilBERT fell to 63.64% (a 35.40-percentage-point drop) a gap of only 0.36 percentage points, equivalent to a single email in the 275-sample adversarial test set. LIME, SHAP, and attention-rollout analysis indicate the two models relied on different evidence yet showed similar vulnerability. Pairwise error analysis shows the models agreed on 54.9% of adversarial samples but each made a similar number of exclusive errors (24 and 25 respectively), indicating partly complementary rather than identical failure modes. The results show that clean-data accuracy does not predict adversarial robustness, and that adversarial testing should be a standard part of phishing-detection evaluation.

发表机构

  • University of Hertfordshire(赫特福德大学)

机构由 AI 辅助整理,请以论文原文为准。

↑