arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越$L_2$:将溯因潜在解释推广到多样化的基于原型的架构

Beyond $L_2$: Generalizing Abductive Latent Explanations to Diverse Prototype-Based Architectures

Jules Soria, Alban Grastien, Romain Xu-Darme, Julien Girard-Satabin, Zakaria Chihani, Daniela Cancila

arXiv 2608.16773首次发表:更新:

发表机构

Université Paris-Saclay; CEA(巴黎-萨克雷大学; 法国原子能和替代能源委员会)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文将溯因潜在解释(ALE)框架推广到支持非欧几里得原型架构,推导了适配不同几何变体的边界算法,通过图像分类器验证了理论构造,实现了对多样化原型架构可解释性的严格跨架构比较。

AI 中文摘要

基于原型的神经网络被称为设计层面可解释的架构。最近,溯因潜在解释(Abductive Latent Explanations,ALE)被提出,以提供形式化、数学上有保证的解释,这些解释利用了这些网络的内在结构,确保了预测安全性和人类可读性。ALE依赖于计算潜在空间距离的紧界来生成形式化解释。然而,现有的ALE公式严格局限于欧几里得潜在空间。这留下了一个关键缺口:现代最先进的架构越来越依赖于非欧几里得表示,例如球形度量、高斯密度和维度投影,这使得当前的形式化解释方法不兼容。在这项工作中,我们将ALE框架推广为支持非欧几里得原型架构。对于每种几何变体,我们系统地推导了如何将架构映射到现有界或构建新颖的、特定于架构的边界算法。我们通过在完全训练好的图像分类器上计算子集极小形式化解释来验证我们的理论构造。通过将这些多样化模型统一在单个形式化框架下,我们首次实现了对它们可解释性的严格跨架构比较。

英文摘要

Prototype-based neural networks are hailed as interpretable-by-design architectures. Recently, Abductive Latent Explanations (ALE) were introduced to provide formal, mathematically guaranteed explanations that leverage the intrinsic structure of these networks to ensure both predictive safety and human readability. ALEs rely on computing tight bounds on latent space distances to produce formal explanations. However, existing ALE formulations are rigidly confined to Euclidean latent spaces. This leaves a critical gap: modern state-of-the-art architectures increasingly rely on non-Euclidean representations - such as spherical metrics, Gaussian densities, and dimensional projections - rendering current formal explanation methods incompatible. In this work, we generalize the ALE framework to support non-Euclidean prototype architectures. For each geometric variant, we systematically derive how to either map the architecture to existing bounds or construct novel, architecture-specific bounding algorithms. We validate our theoretical constructions by computing subset-minimal formal explanations on fully trained image classifiers. By unifying these diverse models under a single formal framework, we enable the first rigorous, cross-architecture comparison of their interpretability.

CommentsAccepted at ECML-PKDD 2026, Research Track

DOI:10.1007/978-3-032-37657-2_7

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑