发表机构
Nanyang Technological University; University of Bern; Griffith University; Xinjiang University(南洋理工大学; 伯尔尼大学; 格里菲斯大学; 新疆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对词元级文本异常检测,提出双证据融合与聚合框架DiFA,结合结构异常与语义不一致性,通过校准融合及多变量聚合实现顶尖性能。
AI 中文摘要
文本异常检测,即识别偏离正常语言模式的文本实例的任务,对于语言驱动的应用至关重要。然而,大多数现有方法只能执行文档级异常检测,这使得难以定位有害短语或支持有针对性的预防。近期,词元级文本异常检测出现了新兴趋势,旨在通过识别文档中的异常单词或片段来解决上述局限。然而,一种代表性方法主要依赖表示空间距离度量,忽略了不同异常线索在捕捉多样异常模式中的互补作用。为弥补这些差距,我们提出了一个具有自适应融合与聚合的双证据框架(DiFA),用于词元级异常检测。DiFA从形式结构和语义两个视角推导异常分数,分别捕捉可见的结构异常和上下文不一致性,从而为识别多样异常提供互补证据。为结合这两个数值尺度不同的分数,DiFA引入校准与融合机制,自适应地平衡两个视角。此外,为获得有判别力的文档级分数,设计了一种多变量聚合方法,从多个角度汇总词元级异常分数,防止稀有异常词元被稀释。在多种文本异常检测基准上的大量实验表明,DiFA持续达到顶尖性能,同时保持强效率、鲁棒性和可解释性。代码和脚本可在以下网址获取:此 https URL。
英文摘要
Text anomaly detection, the task of identifying text instances that deviate from normal language patterns, is crucial for language-driven applications. However, most existing methods can only perform document-level anomaly detection, making it hard to locate harmful phrases or support targeted prevention. Recently, there has been an emerging trend toward token-level text anomaly detection, which aims to address the above limitation by identifying anomalous words or fragments within a document. Nevertheless, one representative method mainly relies on representation-space distance measurement, neglecting the complementary roles of different anomaly cues in capturing diverse abnormal patterns. To bridge the gaps, we propose a Dual-evidence framework with adaptive Fusion and Aggregation (DiFA) for token-level anomaly detection. DiFA derives anomaly scores from form-structural and semantic views to capture visible structural abnormality and contextual inconsistency, respectively, thereby providing complementary evidence for identifying diverse anomalies. To combine these two scores with varying numerical scales, DiFA incorporates a calibration and fusion mechanism to adaptively balance the two views. Moreover, to obtain a discriminative document-level score, a multivariate aggregation method is designed to summarize token-level anomaly scores from multiple perspectives, preventing rare anomalous tokens from being diluted. Extensive experiments across various text anomaly detection benchmarks demonstrate that DiFA consistently achieves top performance while maintaining strong efficiency, robustness, and interpretability. The code and scripts are available at: https://github.com/qyy11-com/DiFA.
CommentsAccepted by ICDM 2026. 10 pages, 4 figures, 2 tables