SIM:基于子空间交互的令牌级文本异常检测方法
SIM: Subspace Interaction-based Method for Token-Level Text Anomaly Detection
浏览论文内容
中文总结 AI 辅助
针对令牌级文本异常检测中局部信号稀释和过度平滑问题,提出基于子空间交互的SIM方法,通过子空间解耦、硬伪异常生成和概率边界损失提升检测性能,实验验证其高效、鲁棒且可解释。
中文摘要 AI 辅助
令牌级文本异常检测作为文本异常检测的新兴趋势,超越了粗粒度的文档级检测,通过定位文本中的异常令牌来实现细粒度异常预测。令牌级文本异常检测在垃圾邮件过滤和虚假新闻检测等众多实际应用中发挥着关键作用。然而,现有方法仍依赖全局距离计算进行评分,在此过程中,局部异常信号被大量冗余的正常特征维度严重稀释。此外,这些方法使用的预训练语言模型不可避免地平滑了表面异常,进一步限制了其在令牌级异常检测中的有效性。为解决这些局限,我们提出了一种基于子空间交互的方法(简称SIM)用于令牌级文本异常检测。为防止局部信号稀释,SIM采用基于子空间交互的异常检测器,将高维令牌嵌入解耦为多个低维嵌入,从而放大隐藏在特定维度中的局部异常信号。为抵消过度平滑效应,我们设计了一个硬伪异常生成模块来构造伪异常令牌,模拟被语义平滑掩盖的细微异常。此外,我们开发了一种概率边界损失,将异常分数标准化为统计距离,有效促使异常实例显著偏离正常分布中心。在多个基准数据集上的大量实验验证了SIM的有效性,并展示了其显著的效率、鲁棒性和可解释性。源代码可在以下网址获取:此https URL。
英文摘要
Token-level text anomaly detection, as an emerging trend of text anomaly detection, moves beyond coarse-grained document-level detection by localizing anomalous tokens within text. By providing fine-grained abnormality prediction, token-level text anomaly detection plays a critical role in various real-world applications, such as spam filtering and fake news detection. However, existing methods still rely on the global distance calculation for scoring, during which the local anomaly signals are severely diluted by numerous redundant normal feature dimensions. Moreover, pre-trained language models used in these methods inevitably smooth out surface anomalies, further limiting their effectiveness in token-level anomaly detection. To address these limitations, we propose a Subspace Interaction-based Method (SIM for short) for token-level text anomaly detection. To prevent local signal dilution, SIM adopts a subspace interaction-based anomaly detector, which decouples high-dimensional token embeddings into multiple low-dimensional ones, amplifying localized anomaly signals hidden within specific dimensions. To counteract the over-smoothing effect, we design a hard pseudo-anomaly generation module to construct pseudo-anomalous tokens, simulating the subtle anomalies obscured by semantic smoothing. Also, a probabilistic boundary loss is developed to standardize anomaly scores into statistical distances, effectively enforcing anomalous instances to deviate significantly from the normal distribution center. Extensive experiments on multiple benchmark datasets verify the effectiveness of SIM and demonstrate its remarkable efficiency, robustness, and interpretability. The source code is available at: https://github.com/yankehan/SIM-TAD.
发表机构
- Guangxi University(广西大学)
- Griffith University(格里菲斯大学)
机构由 AI 辅助整理,请以论文原文为准。