arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

定位全局差异:边际贡献与上下文异常检测

Localizing Global Discrepancies: Marginal Contributions and Contextual Anomaly Detection

Tommaso dorigo

arXiv 2608.28375首次发表:更新:

AI 中文总结

该研究提出框架将重采样诊断等与事件级异常检测关联,推导不同统计量的定位规则,在LHC Olympics基准上验证成对估计量的高效性,还揭示上下文可提供跨事件依赖的额外类别信息。

AI 中文摘要

全局拟合优度和差异统计量可确定样本偏离参考分布,但无法识别哪些观测值导致了这种偏离。我们开发了一个框架来解决该定位问题,通过为每个观测值分配其在随机统计上下文中的条件或边际贡献,将重采样诊断、数据估值与投影理论、事件级异常检测联系起来。对于对称统计量,固定大小替换与中心化条件定位完全等价;对于U统计量,添加得分等于首个Hoeffding/Hájek贡献;对于平滑分布泛函,其主导阶与影响函数相关;对于具有已知背景的无偏MMD,它完全退化为MMD见证函数。该视角还能生成更高效的估计量:匹配上下文减法可去除与观测值无关的波动,而对于成对MMD,含事件项可构成简单定位器。在LHC Olympics异常检测基准上,成对估计量以预测的1/(Rm²)缩放收敛至直接经验MMD见证函数,其中m为批次大小,R为批次数量;当m=1000、R=5×10⁶时,其与直接经验MMD见证函数的相关系数达0.9993,AUC基本完全一致。我们还探究上下文何时包含超出事件自身特征的信息:在共享隐变量玩具模型中,构建的单个事件信号与背景分布完全相同,使得孤立事件AUC=0.5,区分信息仅存在于共享隐变量参数诱导的跨事件依赖中;集成方法可恢复该信息,而独立隐变量对照则无法实现。这区分了上下文的两种作用:一是高效定位全局差异,二是当替代方案包含共享结构时提供真正额外的类别信息。

英文摘要

Global goodness-of-fit and discrepancy statistics can establish that a sample departs from a reference distribution without identifying which observations drive the departure. We develop a framework for this localization problem by assigning to each observation its conditional or marginal contribution across random statistical contexts. This connects resampling diagnostics and data valuation to projection theory and event-level anomaly detection. For symmetric statistics, fixed-size replacement is exactly equivalent to centered conditional localization. For U-statistics, the addition score equals the first Hoeffding/Hájek contribution; for smooth distributional functionals it is related at leading order to the influence function; and for unbiased known-background MMD it reduces exactly to the MMD witness. This viewpoint also yields more efficient estimators. Matched-context subtraction removes fluctuations unrelated to the observation, while for pairwise MMD the event-containing terms give a simple localizer. On the LHC Olympics anomaly-detection benchmark, the pair estimator converges to the direct empirical MMD witness with the predicted 1/(Rm^2) scaling, where m is batch size and R the number of batches. At m=1000 and R=5x106 it reaches correlation 0.9993 with essentially identical AUC. We also ask when context contains information beyond an event's own features. In a shared-latent toy model, the full single-event signal and background distributions are identical by construction, forcing isolated-event AUC=0.5. Discriminating information survives only in cross-event dependence induced by the shared latent parameter; the ensemble recovers this information, whereas an independent-latent control does not. This separates two roles of context: efficient localization of a global discrepancy and genuinely additional class information when the alternative contains shared structure.

Comments32 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑