arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

钓鱼短信检测中的对抗鲁棒性:经典与基于Transformer检测系统的对抗脆弱性对比分析

Adversarial Robustness in Smishing Detection: A Comparative Analysis of Adversarial Fragility in Classical vs. Transformer-Based Detection Systems

Denzel Chiuseni, Athanase Bahizire, Silva Hama, Jema David Ndibwile

arXiv 2608.12889首次发表:更新:

AI 中文总结

本研究对比Random Forest等三类经典词汇模型与mBERT、XLM-RoBERTa两类多语言Transformer的钓鱼短信检测对抗鲁棒性,发现Transformer鲁棒性显著优于经典模型,且干净文本性能无法预测对抗鲁棒性,需针对架构设计防御。

AI 中文摘要

钓鱼短信(smishing)检测系统通常在干净的单语文本上进行训练和评估,但在低资源场景中,攻击者常通过字符混淆、跨语码切换和结构扰动来规避这些系统。本研究针对五种模型架构评估对抗鲁棒性:三种经典词汇模型(Random Forest、XGBoost、CNN+BiLSTM)和两种多语言Transformer(mBERT、XLM-RoBERTa),使用包含27037条消息的数据集。经典模型接受黑盒通用攻击,Transformer则通过注意力引导的定向攻击进行评估,每种模型在三种攻击类型和强度水平下接受测试,以鲁棒性退化率(RDR)衡量性能。结果显示存在明显的架构边界:经典模型在字符混淆和结构扰动下近乎灾难性失效(RDR最高达0.988),而Transformer表现出显著更强的鲁棒性(RDR最高达0.351),其中结构扰动是其最显著的弱点。效应量分析(Cliff's d)表明两类模型间存在显著差异;在Transformer组内,XLM-RoBERTa尽管在干净文本基准上表现更优,但其退化程度高于mBERT。这些发现表明干净文本性能并非对抗鲁棒性的可靠预测指标,采用Mann-Whitney U和Friedman检验的统计验证证实,这些模式源于模型架构而非采样。结果强调了针对特定架构防御的必要性,并将钓鱼短信检测界定为对抗性网络安全挑战,而非静态分类任务。

英文摘要

Smishing detection systems are commonly trained and evaluated on clean, monolingual text. In low-resource settings, however, attackers frequently circumvent these systems through character obfuscation, cross-lingual code-switching, and structural perturbation. This study evaluates adversarial robustness for five model architectures: three classical lexical models (Random Forest, XGBoost, CNN+BiLSTM) and two multilingual transformers (mBERT, XLM-RoBERTa), using a dataset of 27,037 messages. Classical models are subjected to black-box generic attacks, while transformers are evaluated with attention-guided targeting. Each model is tested across three attack types and intensity levels, with performance measured by the Robustness Degradation Ratio (RDR). The results reveal a distinct architectural boundary: classical models experience near-catastrophic failure under character obfuscation and structural perturbation (RDR up to 0.988), whereas transformers demonstrate significantly greater resilience (RDR up to 0.351), with structural perturbation representing their most pronounced vulnerability. Effect-size analysis (Cliff's d) indicates a substantial difference between the two model categories. Within the transformer group, XLM-RoBERTa, despite achieving a higher clean-text baseline, exhibits greater degradation than mBERT. These findings demonstrate that clean-text performance is not a reliable predictor of adversarial robustness. Statistical validation using Mann-Whitney U and Friedman tests confirms that these patterns are attributable to model architecture rather than sampling. The results underscore the necessity for architecture-specific defences and frame smishing detection as an adversarial cybersecurity challenge rather than a static classification task.

Comments14 Pages, 1 Figure, 6 Equations, 4 Tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑