arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

区分求解器,而非方程:用于块隐式模拟的反向扫描伴随方法

Differentiate the Solver, Not the Equation: Reverse-Sweep Adjoints for Block Implicit Simulation

Lei Shu, Ying Jiang, Kui Wu, Yin Yang, Leonidas Guibas, Chenfanfu Jiang

arXiv 2608.08559首次发表:更新:

发表机构

Stanford University; University of California, Los Angeles; LightSpeed Studios, Tencent America; The University of Utah(斯坦福大学; 加州大学洛杉矶分校; 腾讯美国光速工作室; 犹他大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出求解器级微分方法,针对块隐式模拟设计反向扫描伴随求解器,在精度匹配的前提下,比现有方法更快、内存占用更少,可扩展至大规模可微弹性动力学模拟。

AI 中文摘要

可微模拟是学习、控制和逆问题中的关键组件,需要通过非线性隐式求解器的梯度。现有方法要么依赖展开式自动微分,其内存随求解器深度增长;要么依赖方程级隐式微分,该方法会组装全局雅可比矩阵并求解大型稀疏伴随系统,丢弃了前向求解器的局部性——且是对收敛后的方程而非实际运行的有限计算进行微分。我们提出求解器级微分,即对执行的求解器本身进行微分。当求解器由块隐式更新组成时,其离散伴随通过按反向顺序应用相应的伴随更新获得,得到反向扫描公式,其反向传播与前向求解器镜像。从算子角度看,前向传播通过有序局部求解实现近似逆,反向传播通过反向局部伴随求解实现其转置,不构建全局系统。我们将该思想实例化到Vertex Block Descent上,得到一个可微求解器,其反向彩色高斯-赛德尔扫描完全由局部3×3伴随求解组成。在每个求解器深度,反向传播与通过相同执行前向传播得到的自动微分在机器精度上匹配;方程级伴随在一次扫描后误差为37%;在受控的同代码库、同GPU对比中,其比展开式自动微分快33倍,内存消耗少71倍;且相同构造在投影动力学和基于位置的动力学扩展版本上是精确的。我们在单个GPU上将可微弹性动力学扩展到10^6个接触耦合的软体(800万个顶点)。总体而言,本研究强调求解器结构是高效可微模拟的实用组织原则。

英文摘要

Differentiable simulation is a key component in learning, control, and inverse problems, where gradients through nonlinear implicit solvers are required. Existing approaches either rely on unrolled automatic differentiation, whose memory grows with solver depth, or on equation-level implicit differentiation, which assembles global Jacobians and solves large sparse adjoint systems, discarding the locality of the forward solver -- and differentiating the converged equation rather than the finite computation that actually ran. We propose solver-level differentiation, which differentiates the executed solver itself. When a solver is composed of block implicit updates, its discrete adjoint is obtained by applying the corresponding adjoint updates in reverse order, yielding a reverse-sweep formulation whose backward pass mirrors the forward solver. From an operator perspective, the forward pass realizes an approximate inverse through ordered local solves, and the backward applies its transpose through reverse local adjoint solves, constructing no global system. We instantiate this idea on Vertex Block Descent, yielding a differentiable solver whose reverse colored Gauss-Seidel sweeps are composed entirely of local $3\times 3$ adjoint solves. The backward matches automatic differentiation through the identical executed forward to machine precision at every solver depth, where the equation-level adjoint is off by 37% after one sweep; in a controlled same-codebase, same-GPU comparison it is 33x faster and uses 71x less memory than unrolled automatic differentiation; and the same construction is exact on projective dynamics and extended position-based dynamics. We scale differentiable elastodynamics to $10^6$ contact-coupled soft bodies (8M vertices) on one GPU. Overall, this work highlights solver structure as a practical organizing principle for efficient differentiable simulation.

Comments21 pages, 14 figures, 9 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑