arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38497math.NAcs.NAcs.PL

Null-A模式:广义逆的组合计算

The Mode of Null-A: Compositional Computation of a Generalized Inverse

Barak A. Pearlmutter, Jeffrey Mark Siskind

首次发表
浏览论文内容

中文总结 AI 辅助

提出Null-A模式原像自动微分算法,通过组合计算仿射空间原像,推广逆AD问题,支持非方阵Jacobian,适用于CPU小规模与GPU大规模计算。

中文摘要 AI 辅助

我们提出了一种新算法,用于计算通过特殊形式矩阵 ${J}={J}_{T-1}\cdots{J}_0$ 的乘积的仿射空间的原像:找到最大的输入空间 $\mathbf{X}$,使得 $\mathbf{x}\in\mathbf{X}$ 蕴含 ${J}\mathbf{x}\in \mathbf{Y}$,其中 $\mathbf{Y}$ 是给定的输出仿射空间。这些特殊矩阵出现在自动微分(AD)中,其中描述线性化计算的雅可比矩阵 $J$ 恰好具有这种结构:一系列线性化原始数值运算的乘积。这使得我们能够使用新算法来表述 Null-A 模式原像自动微分,它通过数值计算的雅可比矩阵或雅可比转置来寻找仿射原像。这是求解 ${J}\acute{\mathbf{x}}^{\ast}=\acute{\mathbf{y}}^{\ast}$ 或 ${J}^{T}\grave{\mathbf{y}}^{\ast}=\grave{\mathbf{x}}^{\ast}$ 的逆自动微分问题的推广。关键在于以一种适合高效原像计算的方式表示仿射空间,以组合和准局部的方式,通过一系列矩阵 ${J}_t$ 进行。与先前方法不同,Null-A 原像模式自动微分允许 ${J}_t$ 矩阵为非方阵,对应于计算过程中活跃变量数量膨胀和收缩的计算机程序。当 ${J}$ 为方阵且初始仿射空间为单点时,这可以找到常规逆。但在更一般的情况下,拥有整个仿射空间提供了自由度,可以针对特定问题进行利用。我们将该方法应用于 CPU 上的小规模问题,其中 $J_t$ 是线性化的标量一元或二元数值函数;以及 GPU 上的较大规模问题,其中 $J_t$ 是线性化的聚合数组操作,如卷积和注意力。

英文摘要

We present a novel algorithm for calculating the preimage of an affine space through a product ${J}={J}_{T-1}\cdots{J}_0$ of matrices ${J}_t$ of special form: finding the largest input space $\mathbf{X}$ such that $\mathbf{x}\in\mathbf{X}$ implies ${J}\mathbf{x}\in \mathbf{Y}$, where $\mathbf{Y}$ is a given output affine space. These special matrices arise in AD, where the Jacobians $J$ describing the linearized computation have precisely this structure: the product of a series of linearized primitive numeric operations. This allows us to use the new algorithm to formulate Null-A mode preimage AD, which finds the affine preimage through the Jacobian or Jacobian transpose of a numeric computation. This is a generalization of the inverse AD problem of solving ${J}\acute{\mathbf{x}}^{\ast}=\acute{\mathbf{y}}^{\ast}$ or ${J}^{T}\grave{\mathbf{y}}^{\ast}=\grave{\mathbf{x}}^{\ast}$. The key is to represent affine spaces in a fashion which lends itself to efficient preimage calculation, in a compositional and quasi-local fashion, through a succession of matrices ${J}_t$. Unlike previous methods, Null-A preimage mode AD allows the ${J}_t$ matrices to be non-square, corresponding to a computer program whose number of active variables swells and shrinks during the computation. When ${J}$ is square and the initial affine space is a single point, this finds the conventional inverse. But in the more general case, having the entire affine space provides freedom which can be leveraged in a problem-specific manner. We apply the method to small problems on-CPU where the $J_t$ are linearized scalar unary or binary numeric functions; and to larger problems on-GPU where the $J_t$ are linearized aggregate array operations like convolution and attention.

发表机构

  • Maynooth University(梅努斯大学)
  • Purdue University(普渡大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑