arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

可解释架构的图形化设计

Graphical Design of Interpretable Architectures

Pietro Barbiero

arXiv 2608.18936首次发表:更新:

发表机构

IBM Research(IBM研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出改编自Penrose张量表示法的图形符号,可直观呈现可解释AI架构的全局视图并与PyTorch einsum代码一一对应,还绘制了Steerling-8B的架构图并转换为代码。

AI 中文摘要

设计、实现与比较可解释架构需要一种形式化语言来对其进行表示,而最常用的表示方式存在两方面不足:符号方程无法直观呈现架构的全局视图;概率图模型和流程图无法描述实际的张量操作,从而隐藏关键见解并限制可复现性。为弥合这一差距,我们引入一种改编自Penrose张量表示法的图形符号,用于设计可解释AI架构,该图形符号既提供架构的全局视图,又能与PyTorch einsum代码一一对应。我们首先用该符号描述天生可解释的架构,包括概念瓶颈、稀疏探针、原型网络、神经加性模型及线性模型混合体;随后绘制前沿可解释语言模型Steerling-8B的关键架构组件图,该图可提供架构的全局见解(如表明Steerling是残差模型)、各单独操作的几何解释,还能直接转换为33行PyTorch代码。

英文摘要

Designing, implementing, and comparing interpretable architectures requires a formal language to represent them. The most common representations fall short in one of two ways. Symbolic equations give no global view of an architecture at a glance. Probabilistic graphical models and flowcharts do not describe actual tensor manipulations, thus hiding key insights and limiting reproducibility. To close this gap, we introduce a graphical notation for designing interpretable AI architectures, adapted from Penrose tensor notation. This graphical notation gives a global view of an architecture and maps one to one onto PyTorch einsum code. We first use this notation to describe architectures that are interpretable by construction, including concept bottlenecks, sparse probes, prototype networks, neural additive models, and mixtures of linear models. We then diagram the key architectural components of Steerling-8B, a frontier interpretable language model. The diagram yields global insights into the architecture (e.g., showing that Steerling is a residual model), a geometric interpretation of each individual operation, and a direct translation into 33 lines of PyTorch code.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑