arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PhysSAE:基于稀疏自编码器的机制可解释性

PhysSAE: Mechanistic Interpretability of PINNs with Sparse Autoencoders

Nandita N. Patil, Eshwar R. A., Gajanan V. Honnavar

arXiv 2609.07061首次发表:更新:

发表机构

QuaNad Research Laboratory, PES University; QNu Labs Pvt Ltd; PES University (EC Campus)(PES大学 QuaNad研究实验室; QNu Labs私人有限公司; PES大学(EC校区))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

PhysSAE通过稀疏自编码器识别PINN中的物理特征,并因果验证其作用,证明PINN具有可解释的稀疏结构化表示。

AI 中文摘要

物理信息神经网络(PINNs)将偏微分方程残差嵌入神经网络训练中,但其内部表示仍然不透明:其隐藏层编码了哪些物理特征,或者这些特征是否具有局部因果作用,目前尚不清楚。我们提出了PhysSAE,一种机制可解释性框架,该框架在PINN倒数第二层激活上训练过完备稀疏自编码器(SAEs),并通过在原始冻结隐藏状态中直接进行因果干预来评估字典原子:$h_{\mathrm{cf}} = h - \alpha z_k d_k$,完全绕过SAE解码器。在六个PDE族中,每个使用3个PINN种子和3个SAE种子,我们表明:(i)我们发现的SAE原子与独立定义的物理可观测量对齐(最大Pearson $|r|=0.951$,始终$\gg$置换零假设);(ii)顶部对齐原子消融的因果足迹在空间上比PCA或ICA干预更集中1.2-4.2倍;(iii)对于结构化物理概念,顶部对齐原子在因果定位上优于匹配的随机对照(ESF$_{80}$优势0.04-0.44)。双原子双侧表示将概念回归R$^2$提高了$\Delta R^2\\!=\\!0.05\text{-}0.15$,而随机对则将其降低最多0.60。这些结果表明,PINNs发展出稀疏、物理结构化的潜在表示,可以在事后被识别和因果询问,为可解释性感知的科学机器学习开辟了道路。

英文摘要

Physics-Informed Neural Networks (PINNs) embed PDE residuals into neural network training, but their internal representations remain opaque: it is unknown what physical features their hidden layers encode or whether those features have a localized causal role. We present PhysSAE, a mechanistic interpretability framework that trains overcomplete sparse autoencoders (SAEs) on PINN penultimate-layer activations and evaluates dictionary atoms through direct causal intervention in the original frozen hidden state: $h_{\mathrm{cf}} = h - αz_k d_k$, bypassing the SAE decoder entirely. Across six PDE families, with 3 PINN seeds and 3 SAE seeds each---we show that (i) Our discovered SAE atoms align with independently-defined physical observables (max Pearson $|r|=0.951$, always $\gg$ permutation null), (ii) the causal footprint of top-aligned atom ablation is 1.2--4.2$\times$ more spatially concentrated canonical than PCA or ICA interventions, and (iii) top-aligned atoms outperform matched random controls on causal localization for structured physical concepts (ESF$_{80}$ advantage 0.04-0.44). Two-atom bilateral representations improve concept regression R$^2$ by $ΔR^2\!=\!0.05\text{-}0.15$ over single atoms, while random pairs decrease it by up to 0.60. These results demonstrate that PINNs develop sparse, physically structured latent representations that can be identified and causally interrogated post-hoc, opening a path toward interpretability-aware scientific machine learning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑