发表机构
Università degli Studi di Milano; INFN Sezione di Milano; Niels Bohr Institute, University of Copenhagen(米兰大学; 米兰国家核物理研究所; 哥本哈根大学尼尔斯·玻尔研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究将基于稀疏自编码器的机械可解释性方法应用于中微子基础模型,识别出经验证的物理概念图谱,训练的不确定性预测头可提升20%选择效率下的中位角分辨率,证明机械可解释性可助力下游任务设计。
AI 中文摘要
我们首次将基于稀疏自编码器的机械可解释性方法应用于粒子物理学。研究了在IceCube数据上预训练并针对方向重建进行微调的中微子基础模型,我们采用包含保留测试集、匹配的干扰项控制及独立字典训练复现的严格验证方案,在模型表示中识别出经验证的物理概念图谱。因果干预显示,方向预测头几乎未利用该图谱。受此未充分利用的信息启发,我们在同一事件级表示上训练不确定性预测头,以预测模型的角重建误差。与方向预测头不同,该不确定性预测头因果依赖于图谱中的质量和亮度特征。在20%的选择效率下,此可解释估计器将中位角分辨率从20.2°提升至3.2°。这些结果表明,机械可解释性可揭示模型内部表示中编码的学习到的隐式物理规律,并助力设计利用该规律的下游任务。
英文摘要
We present a first application of sparse-autoencoder-based mechanistic interpretability to particle physics. Studying a neutrino foundation model pretrained on IceCube data and fine-tuned for direction reconstruction, we identify a validated atlas of physical concepts in the model representation, using a strict validation protocol consisting of held-out tests, matched nuisance controls, and replication across independent dictionary trainings. Causal interventions show that the direction head barely draws on this atlas. Motivated by this underused information, we train an uncertainty head on the same event-level representation to predict the model's angular reconstruction error. Unlike the direction head, it depends causally on quality and brightness features from the atlas. At $20\%$ selection efficiency, this interpretable estimator improves the median angular resolution from $20.2^\circ$ to $3.2^\circ$. These results suggest that mechanistic interpretability can reveal learned latent physics encoded within a model's internal representation and help design downstream tasks that exploit it.