发表机构
New York University(纽约大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
RareSense是面向交易数据异常检索的稀有度感知相似度框架,通过挖掘稀有项集与关联规则实现相似度计算,在多领域基准实验中其检索性能优于经典度量且与专用检测器相当。
AI 中文摘要
针对稀疏集合型数据的相似度搜索常受频繁背景属性主导,因为Jaccard、余弦、汉明等经典度量通过原子重叠比较对象。IDF(逆文档频率)加权可部分缓解该影响,但仍基于原子层面,无法显式表征有价值的高阶共现关系。本文提出RareSense,一种面向稀疏交易型异常数据的稀有度感知相似度框架。RareSense挖掘最小稀有项集作为中间结构,推导可靠的稀有关联规则,将对象映射为稀疏稀有规则特征,再用加权Jaccard相似度比较。规则权重结合逆支持度、置信度、提升度、结构复杂度和稳定性,使邻域由共享的稀有证据而非均匀特征重叠决定。研究表明,IDF加权Jaccard是RareSense的受限单元素特例,诱导的距离在原始对象上为伪度量,在相同规则特征定义的等价类上为度量。在涵盖网络安全和通用分类领域的四个基准系列实验中,RareSense在所有评估的相似度度量中达到最高的查询条件检索宏平均性能。统计分析显示存在显著整体差异,经校正的配对比较表明RareSense优于原子级基线。性能增益仍依赖工作负载,当异常共享可重复的稀有高阶结构时增益最强。对于全局异常排名,RareSense达到最高的宏平均性能,同时与多个强大的专用检测器在统计上具有可比性。
英文摘要
Similarity search over sparse set-valued data is often dominated by frequent background attributes because classical measures such as Jaccard, cosine, and Hamming compare objects through atomic overlap. IDF (Inverse document frequency) weighting partially reduces this effect but remains atom-wise and cannot explicitly represent informative higher-order co-occurrences. We introduce RareSense, a rarity-aware similarity framework for sparse transactional anomaly data. RareSense mines minimal rare itemsets as intermediate structures, derives reliable rare association rules, maps objects into sparse rare-rule profiles, and compares them using weighted Jaccard similarity. Rule weights combine inverse support, confidence, lift, structural complexity, and stability, so that neighborhoods are determined by shared rare evidence rather than uniform feature overlap. We show that IDF-weighted Jaccard is a restricted singleton case of RareSense, and that the induced distance is a pseudometric on the original objects and a metric over equivalence classes defined by identical rule profiles. Experiments across four benchmark families spanning cybersecurity and general categorical domains show that RareSense attains the highest observed macro-average query-conditioned retrieval performance among the evaluated similarity measures. The statistical analysis indicates significant overall differences, with corrected paired comparisons favoring RareSense over the atomic baselines. The gains remain workload-dependent and are strongest when anomalies share repeatable rare higher-order structure. For global anomaly ranking, RareSense achieves the highest observed macro-average performance while remaining statistically comparable to several strong dedicated detectors.