AI 中文总结
本文提出X-WAD方法,采用基于Transformer的语言模型,通过token级logit的惊讶度映射为HTTP请求异常检测提供可解释性,还发现公开数据集的标签不一致问题。
AI 中文摘要
基于Web的服务(尤其是API驱动架构)的快速发展,反映出对分布式系统的依赖日益加深,这使得敏感数据面临安全风险,因此采用自动防御机制至关重要。在实际场景中良性流量占主导的情况下,现代防御手段越来越多地对正常行为进行建模,依赖仅用正常数据训练的半监督方法。但在实践中很难确保此类训练数据中完全没有异常实例,而标签错误或被污染的攻击样本会给学习到的防御机制引入后门,导致模型将某些攻击模式错误分类为正常行为。本文研究基于Transformer的语言模型(TLM)在HTTP请求异常检测中的有效性,重点为检测到的异常提供详细解释。该研究采用基于token级logit的惊讶度映射,同时提供异常分数和类似热力图高亮的直接详细解释。通过在一个流行的公开数据集中发现标签不一致,证明了所提出可解释性方法的有效性,揭示了训练数据中的异常污染如何在检测模型中诱导类似后门的失效。
英文摘要
The rapid growth of web-based services, particularly API-driven architectures, reflects an increasing reliance on distributed systems, exposing sensitive data to security risks and making the adoption of automated defensive mechanisms essential. In this context, where benign traffic predominates in real-world settings, modern defenses increasingly model normal behavior, relying on semi-supervised approaches trained on only normal data. However, ensuring the complete absence of anomalous instances in such training data is inherently difficult in practice, and mislabeled or contaminated attack samples can introduce backdoors into the learned defense, causing the model to silently misclassify certain attack patterns as normal behavior. This paper investigates the effectiveness of Transformer-based Language Models (TLMs) in the detection of anomalies in HTTP requests, focussing on providing detailed explanations for the detected anomalies. The study employs token-level logit-based surprisal mapping to provide both an anomaly score and a direct, detailed explanation via heatmap-like highlighting. The effectiveness of the proposed explainability approach is demonstrated by the discovery of labelling inconsistencies in a popular public dataset, revealing how anomalous contamination in the training data had induced backdoor-like failures in the detection models.