arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.01171cs.CR

Johnny仍收到垃圾短信:评估SMS垃圾短信检测的鲁棒性

Johnny Still Receives Spam SMS: Assessing the Robustness of SMS Spam Detection

Muhammad Salman, Muhammad Islam, Muhammad Ikram, Mohamed Ali Kaafar

首次发表
浏览论文内容

中文总结 AI 辅助

本文评估了终端用户依赖的SMS反垃圾短信系统的鲁棒性,发现其对抗编码级攻击的鲁棒性不足,提出多模型集成方法可大幅提升对抗攻击的鲁棒性并保持分类准确率。

中文摘要 AI 辅助

SMS垃圾短信检测系统在受控环境中常能达到高准确率,但在实际部署中面临对抗性攻击和日益复杂的垃圾短信策略时表现不佳。本文评估了终端用户实际依赖的SMS反垃圾短信系统的鲁棒性,包括商业消息应用、第三方反垃圾短信服务以及Hugging Face上托管的公开开放权重模型。我们在标准和对抗条件下对这些系统进行评估,考虑了可感知攻击和最先进的不可感知攻击,仅纳入经我们验证可在真实SMS或RCS传输中存活的扰动,而非仅存在于实验室的人工产物。实验表明,现有垃圾短信检测器识别对抗性操纵消息的能力存在显著差距;我们进一步证明仅靠对抗性训练是不够的。采用明确的保留评估协议,我们发现鲁棒性在同一扰动族内的迁移效果良好,但在结构不同的编码级攻击下会急剧下降。为解决这些弱点,我们提出一种多模型集成方法,将对抗性训练与架构和分词方式多样的垃圾短信分类器相结合。结果显示,该集成方法尤其在采用少数投票策略时,可大幅提升对可感知和不可感知对抗性攻击的鲁棒性,同时保持有竞争力的分类准确率;我们还分析了所得的精确率-召回率权衡关系,并为对误报敏感和对召回率关键的部署场景推荐了操作点。这些发现凸显了在实际场景中构建更安全的SMS垃圾短信检测系统,需要开展全面的鲁棒性评估并采用基于集成的防御方法。

英文摘要

SMS spam detection systems often achieve high accuracy in controlled environments but struggle against adversarial attacks and increasingly sophisticated spam tactics in real-world deployments. In this paper, we evaluate the robustness of SMS anti-spam systems that end users actually rely on, including commercial messaging applications, third-party anti-spam services, and publicly available open-weight models hosted on Hugging Face. We evaluate these systems under both standard and adversarial conditions, considering perceptible and state-of-the-art imperceptible attacks. We include only perturbations that we verify survive real SMS or RCS delivery, rather than lab-only artifacts. Our experiments reveal significant gaps in existing spam detectors' ability to identify adversarially manipulated messages. We further demonstrate that adversarial training alone is insufficient. Using an explicit held-out evaluation protocol, we find that robustness transfers well within a perturbation family but degrades sharply against structurally distinct, encoding-level attacks. To address these weaknesses, we propose a multi-model ensemble that combines adversarial training with spam classifiers diverse in architecture and tokenization. Our results show that this ensemble, particularly when using a minority-voting strategy, substantially improves robustness against both perceptible and imperceptible adversarial attacks while maintaining competitive classification accuracy. We also characterize the resulting precision-recall trade-off and recommend operating points for false-positive-sensitive and recall-critical deployments. These findings highlight the need for comprehensive robustness evaluations and ensemble-based defenses for building more secure SMS spam detection systems in real-world settings.

发表机构

  • Macquarie University(麦考瑞大学)
  • University of Engineering and Technology Mardan(曼德工程与技术大学)

机构由 AI 辅助整理,请以论文原文为准。

↑