arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.13652cs.LGhep-exhep-ph

对撞机实验中可解释异常检测的对比学习

Contrastive Learning for Interpretable Anomaly Detection at Collider Experiments

Haoyi Jia, Sagar Addepalli, Julia Gonski

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对对撞机异常检测的可解释性与能量依赖问题,提出ORCA两阶段框架,通过对比学习结合自编码器实现更优的信号灵敏度与可解释性,为对撞机异常搜索提供新途径。

中文摘要 AI 辅助

对撞机物理中通用的事例级异常检测存在两个反复出现的问题:异常分数难以解释,且与能量标度和粒子多重数强相关。我们提出了基于对比学习的异常检测有序表示框架(Organized Representation via Contrastive learning for Anomaly detection,ORCA),这是一个两阶段框架:首先通过对多样化物理过程的监督对比学习学习嵌入空间,随后在该空间中运行标准自编码器以生成事例级异常分数。在与高亮度大型强子对撞机(High-Luminosity Large Hadron Collider)条件一致的模拟数据集上,相较于基线自编码器架构,ORCA在对新物理信号的灵敏度广度和深度上均取得显著提升。除灵敏度提升外,对比嵌入还使异常样本具备可解释性:由于已知过程占据嵌入空间的不同区域,对嵌入分布进行最大似然模板拟合可将异常样本中的事例归因于具有量化不确定性的模板物理过程。我们证明该拟合能准确恢复注入的信号产额,包括嵌入训练中未包含的信号,还能通过异常样本最相似的已知过程来表征模板库中不存在的信号。这些结果确立了ORCA作为对撞机基于异常检测的可解释搜索的一种途径,其嵌入几何结构相比标准一维输出拟合承载了更高维度的物理信息,增强了下游统计分析。

英文摘要

Generic event-level anomaly detection for collider physics has two recurring problems: anomaly scores are hard to interpret, and they correlate strongly with energy scale and object multiplicity. We present Organized Representation via Contrastive learning for Anomaly detection (ORCA), a two-stage framework that first learns an embedding space via supervised contrastive learning across a diverse set of physics processes, then runs a standard autoencoder in that space to generate event-level anomaly scores. On a simulated dataset consistent with conditions at the High-Luminosity Large Hadron Collider, ORCA delivers significant gains in both breadth and depth of sensitivity to new physics signals with respect to a baseline autoencoder architecture. Beyond improved sensitivity, the contrastive embedding makes the anomalous sample interpretable: because known processes occupy distinct regions of the space, a maximum-likelihood template fit to the embedding distributions can attribute events in an anomalous sample to template physics processes with quantified uncertainties. We demonstrate that the fit accurately recovers injected signal yields, including for signals excluded from the training of the embedding, and characterizes signals absent from the template library through the known processes they most resemble. These results establish ORCA as a route to interpretable anomaly detection-based searches at colliders, where the embedding geometry carries higher dimensional physics information compared to standard one-dimensional output fits, enhancing downstream statistical analysis.

发表机构

  • Stanford University(斯坦福大学)
  • SLAC National Accelerator Laboratory(SLAC国家加速器实验室)

机构由 AI 辅助整理,请以论文原文为准。

↑