发表机构
BIFOLD – Berlin Institute for the Foundations of Learning and Data; LVMT, ENPC, Institut Polytechnique de Paris, Univ Gustave Eiffel; Machine Learning Group, Technical University of Berlin(柏林学习与数据基础研究所; 法国国立路桥学校,巴黎理工学院,古斯塔夫·埃菲尔大学; 柏林工业大学机器学习组)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
asdex是JAX生态系统中首个自动稀疏微分工具包,通过四步流程利用导数矩阵的稀疏性,减少AD过程数量,为相关组件提供稀疏替代实现。
AI 中文摘要
科学计算与机器学习中的许多任务都需要函数的雅可比矩阵或黑塞矩阵。自动微分(AD)可计算这些导数至机器精度,但要生成稠密的m×n雅可比矩阵,需执行n次前向模式AD或m次反向模式AD,每次对应一列或一行。对于一大类函数,每个输出仅依赖少数输入,这使得导数矩阵是稀疏的。自动稀疏微分(ASD)通过四个步骤利用该结构:检测与输入无关的稀疏模式、对图进行着色以分组可共享一次AD过程的列或行、压缩微分以每种颜色执行一次AD过程来计算压缩导数矩阵,最后解压为原始稀疏模式。颜色数量(即AD过程的数量)通常与问题维度无关:例如,具有b个连续带的带状雅可比矩阵,无论其规模多大,仅需b种颜色。asdex是流行JAX生态系统中首个独立的ASD工具包,提供了相关链接,为对应组件提供了稀疏的替代实现。
英文摘要
Many tasks in scientific computing and machine learning require the Jacobian or Hessian matrix of a function. Automatic differentiation (AD) computes these derivatives to machine precision, but materializing a dense $m \times n$ Jacobian requires $n$ forward-mode or $m$ reverse-mode AD passes, one per column or row. For a large class of functions, each output depends on only a few inputs, making the derivative matrix sparse. Automatic sparse differentiation (ASD) exploits this structure in four steps: detection of the input-agnostic sparsity pattern, coloring of a graph to group columns or rows that can share an AD pass, compressed differentiation to compute a compressed derivative matrix with one AD pass per color, and finally decompression into the original sparsity pattern. The number of colors, and hence of AD passes, is often independent of the problem dimension: a banded Jacobian with $b$ contiguous bands, for instance, only ever requires $b$ colors, regardless of its size. asdex offers the first standalone ASD toolkit in the popular JAX ecosystem. With asdex.jacobian and asdex.hessian, it provides sparse drop-in replacements for jax.jacobian and jax.hessian.
Comments1 table