发表机构
Federal University of Bahia; Grupo de Estudos e Aplicação de Inteligência Artificial em Geofísica (GAIA)(巴伊亚联邦大学; 地球物理学人工智能研究与应用组(GAIA))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对深度均衡模型训练中伴随系统因雅可比近奇异导致梯度放大的问题,提出响应重归一化框架,通过CMR等方法控制近临界伴随放大,提升参数更新可靠性且保留有效梯度信息。
AI 中文摘要
深度均衡模型(DEQs)从模型更新未改变的隐表示计算预测值。通过该均衡进行训练需使用隐式求导,并要求求解由残差雅可比矩阵构建的伴随系统。若该雅可比矩阵沿损失敏感方向接近奇异,微小扰动会在伴随响应中被强烈放大,产生大且高度敏感的梯度,可能导致优化不可靠。我们提出响应重归一化(Response Renormalization),这是一种反向传播框架,可提升选定的近极点分母,同时保持未提升的响应通道不变。集体模式响应重归一化(CMR)在低维临界子空间应用该校正,而Phi自适应CMR则根据正 susceptibility 规则计算有界响应质量。我们推导了稠密和无矩阵的集体公式,区分了修改后的冻结锚残差的精确梯度与反向响应代理的梯度,并将该构造扩展至结构化隐层和向量吸引子(SILVA)。在涵盖偏微分方程、三维场、算子映射、复杂几何和粒子系统的23个多物理族中,在超过98%的静态族-种子比较和95%的瞬态族-种子比较中,CMR和Phi-CMR的测试误差比使用精确隐式求导训练的模型的测试误差高不超过5%。求解器索引实验显示其收敛至静态伴随,而物理时间滚动在评估条件下保持预测保真度。这些结果表明,选择性响应重归一化可控制近临界伴随放大,同时不全局阻尼良态灵敏度,因此该方法可使参数更新更可靠,同时保留学习所需的有用梯度信息。
英文摘要
Deep Equilibrium Models (DEQs) compute predictions from a hidden representation unchanged by the model update. Training through this equilibrium uses implicit differentiation and requires solving an adjoint system built from the residual Jacobian. If this Jacobian is nearly singular along loss-sensitive directions, small perturbations can be strongly amplified in the adjoint response, producing large, highly sensitive gradients that can make optimization unreliable. We introduce Response Renormalization, a backward-pass framework that lifts selected near-pole denominators while leaving unlifted response channels unchanged. Collective Mode Response Renormalization (CMR) applies this correction in a low-dimensional critical subspace, while Phi-adaptive CMR computes a bounded response mass from a positive susceptibility rule. We derive dense and matrix-free collective formulations, distinguish exact gradients of a modified frozen-anchor residual from backward-response surrogates, and extend the construction to Structured Implicit Layers and Vector Attractors (SILVA). Across 23 multiphysics families spanning partial differential equations, three-dimensional fields, operator maps, complex geometries, and particle systems, CMR and Phi-CMR yield test errors no more than five percent higher than those from models trained with exact implicit differentiation in more than 98% of static and 95% of transient family-seed comparisons. Solver-index experiments show convergence toward the static adjoint, while physical-time rollouts retain predictive fidelity under the evaluated conditions. These results demonstrate that selective response renormalization can control near-critical adjoint amplification without globally damping well-conditioned sensitivity. Therefore, the method can make parameter updates more reliable while preserving the useful gradient information needed for learning.
Comments42 pages, 41 figures