arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用机器学习增强Web应用防火墙以检测SQL注入攻击

Enhancing Web Application Firewalls with Machine Learning for SQL Injection Detection

Lilliane Linnet Musoke, Atta Badii, Ahmed Ashlam

arXiv 2608.28889首次发表:更新:

发表机构

University of Reading(雷丁大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究设计并优化了DistilBERT-堆叠集成模型,提升了SQL注入检测的准确率与鲁棒性,大幅降低推理延迟,为构建实时高效的Web应用防火墙提供了新方案。

AI 中文摘要

检测SQL注入(SQLi)攻击是Web应用安全领域最关键的挑战之一。本研究开展系统文献综述,明确该领域的研究缺口,针对性设计并优化了DistilBERT-堆叠集成管道,以提升检测效率与鲁棒性,同时降低误报率与漏报率。研究执行了全面的预处理与分词操作,提取DistilBERT嵌入,训练机器学习及集成分类器,并基于准确率、精确率、召回率与F1分数对其排序;将表现最佳的三个模型(逻辑回归、XGBoost、SVM)通过神经元学习器组合为堆叠集成模型。该集成模型经快速梯度符号法(FGSM)生成的对抗样本强化,并使用Optuna调优。优化后的集成模型在所有报告指标上达到99.81%,与最强单模型(DistilBERT-SVM,99.82%)表现接近。在本研究使用的评估平台(第3.8节)上,该集成模型对完整测试集的分类耗时为0.0136秒,而DistilBERT-SVM耗时1.896秒,推理延迟降低约140倍,同时在单步FGSM攻击下仍保持99.77%的准确率。本研究的贡献是设计并验证了一款SQLi检测器,其具备与当前最优水平相当的准确率、实时处理速度,且对单步FGSM攻击表现出鲁棒性,而非仅实现准确率的微小提升;敏感性分析进一步确认了模型的稳定性。这些发现凸显了对抗训练与堆叠元学习在构建用于SQLi检测的鲁棒Web应用防火墙(WAF)中的价值。为便于开放验证,数据集、测试集及模型已在指定URL提供。

英文摘要

Detecting SQL Injection (SQLi) attacks ranks among the most critical challenges in web application security. This research conducted a systematic literature review to identify the research gaps in this domain and responsively designed and optimised a DistilBERT-Stacked Ensemble pipeline to improve detection efficiency and robustness while reducing false-positive and false-negative rates. Comprehensive pre-processing and tokenisation were performed, DistilBERT embeddings were extracted, and machine-learning and ensemble classifiers were trained and ranked on accuracy, precision, recall and F1-score. The three best performers (Logistic Regression, XGBoost and SVM) were combined through a neural meta-learner to form a stacked ensemble. The ensemble was hardened with adversarial examples generated by the Fast Gradient Sign Method (FGSM) and tuned with Optuna. The optimised ensemble achieved 99.81% across all reported metrics, closely comparable to the strongest single model (DistilBERT SVM, 99.82%). On the evaluation platform used in this study (Section 3.8), the ensemble classified the full test set in 0.0136s against 1.896s for DistilBERT-SVM, an approximately 140-fold reduction in measured inference latency, while retaining 99.77% accuracy under a single-step FGSM attack. The contribution is the design and validation of a SQLi detector performing with state-of-the-art accuracy at real-time speed and with demonstrated robustness to a single-step FGSM attack, rather than a marginal gain in accuracy. Sensitivity analysis further confirmed the stability of the model. These findings highlight the value of adversarial training and stacked meta-learning in building robust Web Application Firewalls (WAFs) for SQLi detection. For open validation, the dataset, test sets and models are made available at https://github.com/mlily2024/Final-project-SQL-injection-pipeline.

Comments12 pages, 8 figures, 6 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑