arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.28922cs.LG

RankShift:数据库内分类分布偏移的检测与解释

RankShift: In-Database Detection and Explanation of Categorical Shifts

  • Microsoft(微软)

机构由 AI 辅助整理,请以论文原文为准。

Omair Shafi Ahmed

AI总结:

RankShift是一种数据库内分类分布偏移检测方法,无需模型训练,在HDFS、BGL等数据集上检测性能接近或优于计数向量自编码器,可检测罕见类别偏移且误报率符合要求。

AI中文摘要:

登录服务可能收到的失败登录请求数量与往常相同,但某一来源占比从2%升至30%;系统日志中也会出现类似模式:某一罕见事件模板变得常见,但消息速率保持稳定,这类事件会改变活跃的类别,而不改变事件发生的数量。RankShift可在存储数据的分析数据库内检测此类变化,它会将每个窗口的类别占比与良性参考值进行比较,使用皮尔逊分数,其组成部分可识别导致变化的类别。同一查询会返回分数、校准后的警报以及最大的增长贡献项。我们在HDFS、BGL和Thunderbird数据集上对RankShift进行评估,在HDFS上,它的AUROC与计数向量自编码器相差0.001(分别为0.999和1.000),在Thunderbird上表现更优(分别为0.983和0.949)。在受控的固定数量实验中,RankShift可检测到事件计数监控无法察觉的罕见类别偏移,其AUROC达到0.787,而自编码器为0.771。在所有三个语料库上,观测到的误报率与请求的操作水平相符。RankShift无需模型训练或推理服务,而部署的自编码器状态规模是其137倍。

英文摘要:

A login service can receive its usual number of failed sign-ins while one source grows from 2% to 30% of them. The same pattern appears in system logs when a rare event template becomes common while the message rate stays stable. These events change which categories are active without changing how many events occur. RankShift detects such changes inside the analytical database that stores the data. It compares each window's category shares with a benign reference using a Pearson score whose terms identify the categories responsible for the change. The same query returns the score, calibrated alert, and largest increasing contributions. We evaluate RankShift on HDFS, BGL, and Thunderbird. It matches the count-vector autoencoder within 0.001 AUROC on HDFS (0.999 versus 1.000) and leads on Thunderbird (0.983 versus 0.949). In a controlled fixed-volume experiment, RankShift detects rare-category shifts that are invisible to event-count monitoring, reaching 0.787 AUROC compared with 0.771 for the autoencoder. Across all three corpora, observed false-alarm rates track the requested operating levels. RankShift requires no model training or inference service, and the autoencoders deployed state is 137x larger.

↑