AI 中文总结
该研究构建四层语义评估基准PayloadSemBench,系统评估LLM对Web攻击载荷的理解深度,发现严重性评估是主要弱点,且语义理解与检测决策部分解耦。
AI 中文摘要
通过Web接口和API提供的计算机视觉服务处理用于图像资源获取、推理任务配置和结果管理的文本请求,这使得Web攻击载荷分析与其部署安全性密切相关。大型语言模型(LLM)能够识别载荷类型并解释攻击意图。然而,现有研究通常将载荷分析视为单层分类任务,既缺乏对LLM理解载荷深度的系统性评估,也缺乏专门针对Web攻击载荷语义理解深度的评估基准。我们构建了PayloadSemBench,一个四层语义评估基准,将载荷理解操作化为可测量的任务,包含240个载荷。其真实标注通过两轮锚定校准建立,并由第四位独立专家重新验证。两项实验,一项语义理解基准测试和一项将语义理解映射到检测性能的分析,产生了三个主要发现:(1)类型识别和意图理解通常表现强劲,而严重性评估是主要弱点;(2)混淆的影响因模型和层次而异,意图解释和特定混淆技术的重建更容易退化,而性能并非在所有层次上同步下降;(3)语义理解和检测决策部分解耦,只有12.5%至50%的漏检可归因于语义理解失败。在包含生产环境WAF告警流和真实应用请求的独立180条记录数据集上的外部重新评估,重现了非均匀的四层能力概况以及混淆载荷上的层次特定差异。
英文摘要
Computer vision services delivered through Web interfaces and APIs process textual requests for image-resource acquisition, inference-task configuration, and result management, making Web attack-payload analysis relevant to their deployment security. Large language models (LLMs) can identify payload types and explain attack intent. However, existing studies generally treat payload analysis as a single-layer classification task and lack both a systematic assessment of how deeply LLMs understand payloads and an evaluation benchmark dedicated to the depth of semantic understanding of Web attack payloads. We construct PayloadSemBench, a four-layer semantic evaluation benchmark that operationalizes payload understanding across measurable tasks and comprises 240 payloads. Its ground truth was established through two rounds of anchor calibration and re-verified by a fourth independent expert. Two experiments, a semantic-understanding benchmark and an analysis mapping semantic understanding to detection performance, yielded three main findings: (1) type identification and intent understanding were generally strong, whereas severity assessment was the principal weakness; (2) the effects of obfuscation varied across models and layers, with intent explanation and reconstruction of specific obfuscation techniques more susceptible to degradation, while performance did not degrade synchronously across all layers; and (3) semantic understanding and detection decisions were partially decoupled, with only 12.5% to 50% of missed detections attributable to semantic-understanding failures. External re-evaluation on an independent 180-record dataset comprising production WAF alert streams and real application requests reproduced the non-uniform four-layer capability profile and the layer-specific differences on obfuscated payloads.
CommentsAccepted at the 6th International Conference on Computer Vision, Application and Algorithm (CVAA 2026). 16 pages