发表机构
University of Reading(雷丁大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出优化的混合BERT-GNN流水线,经Optuna调优后在SQL注入检测上取得高准确率,鲁棒性良好,相关资源已开放验证。
AI 中文摘要
检测复杂的SQL注入(SQLi)攻击仍是Web应用安全领域最关键的挑战之一。本研究提出了一种优化的混合BERT-GNN流水线,可提升检测准确率与鲁棒性,同时降低误报率与漏报率。该方法将SQL查询分词后编码为上下文BERT嵌入,以此初始化图神经网络(GNN)的节点特征,GNN经训练后对每个查询进行分类,其架构通过Optuna工具针对准确率、精确率、召回率与F1值进行调优。所提模型在攻击类别上取得了99.67%的准确率、99.71%的精确率、99.39%的召回率及99.55%的F1值。通过对图输入施加扰动的敏感性分析进一步评估了模型鲁棒性,得到的平均敏感性得分仅为0.0037,表明模型在这类扰动下预测稳定。研究结果证明了将BERT的上下文理解能力与GNN的结构建模能力相结合的新型混合模型在检测复杂SQLi攻击向量方面的潜力。为便于开放验证,数据集、测试集及模型已在该https地址提供。
英文摘要
Detecting sophisticated SQL Injection (SQLi) attacks remains among the most critical challenges in web applications security. This research study has resulted in an optimised hybrid BERT-GNN pipeline with improved detection accuracy and robustness while reducing false-positive and false-negative rates. SQL queries are tokenised and encoded into contextual BERT embeddings, which then initialise the node features of a Graph Neural Network (GNN) trained to classify each query, with the architecture tuned by Optuna over accuracy, precision, recall, and F1-score. The proposed model achieved 99.67% accuracy, with 99.71% precision, 99.39% recall, and 99.55% F1-score on the attack class. A sensitivity analysis, performed by perturbing graph inputs, further assessed the model robustness and yielded a low mean sensitivity score of 0.0037, indicating stable predictions under such perturbations. The results have demonstrated the potential of a novel hybrid model that couples BERT contextual understanding with the GNN structural modelling to detect sophisticated SQLi attack vectors. For open validation, the dataset, test sets and models are made available at https://github.com/mlily2024/Final-project-SQL-injection-pipeline.
Comments13 pages, 11 figures, 4 tables