发表机构
Aarhus University; Technical University of Denmark(奥胡斯大学; 丹麦技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对复值神经网络训练中步长规则忽略复平面角信息的问题,提出AURA逐参数步长自适应方法,可叠加于任意一阶优化器,通过测量连续更新的一致性动态调整步长,实验表明在多数情况下以低开销提升收敛。
AI 中文摘要
复值神经网络(CVNN)越来越多地被用于处理复值数据;然而,它们通常使用从实值情形继承而来的一阶优化器进行训练。这些方法的效率在很大程度上取决于步长,而它们的步长规则忽略了复平面中可用的角信息。我们通过引入AURA(角更新率自适应)来解决复域中的步长自适应问题,这是一种逐参数的步长自适应方法,可以添加到任何一阶优化器之上,也可以从中移除,而不改变其更新方向。AURA测量每个复参数连续更新之间在长度、对齐和旋转方向方面的一致性,当它们一致时增大步长,不一致时减小步长。它不需要额外的梯度评估,每一步仅需廉价的向量运算。我们将AURA与Adam和Muon结合,并在四个复杂度递增的测试用例(从标量复函数的逼近到物理信息训练)中,将所得方法与著名的一阶优化器进行比较。本工作全程使用全连接神经网络。除步长外的所有超参数在各测试用例中保持固定;对于其中一个用例,我们还在相同预算下对每个优化器的超参数进行了调优。我们的实证测试表明,AURA在大多数情况下以较小的每步开销改善了其基础优化器的收敛性,并且我们确定了它无法做到这一点的条件。
英文摘要
Complex-valued neural networks (CVNNs) are increasingly adopted for complex-valued data; however, they are often trained with first-order optimizers inherited from the real-valued case. The efficiency of these methods depends largely on the step size, and their step-size rules ignore the angular information available in the complex plane. We address step-size adaptation in the complex domain by introducing AURA (Angular Update Rate Adaptation), a per-parameter step-size adaptation that can be added on top of any first-order optimizer, and removed from it, without altering its update direction. AURA measures the agreement between consecutive updates of each complex parameter, in length, alignment, and sense of rotation, and enlarges the step when they are consistent and reduces it when they are not. It requires no additional gradient evaluations and only inexpensive vector operations per step. We combine AURA with Adam and Muon and compare the resulting methods with well-known first-order optimizers on four test cases of increasing complexity, ranging from the approximation of scalar complex functions to physics-informed training. Fully connected neural networks are used throughout this work. All hyperparameters other than the step size are held fixed across test cases; for one case, we also tune the hyperparameters of each optimizer under the same budget. Our empirical tests show that AURA improves the convergence of its base optimizer in most cases with a small per-step overhead, and we identify the conditions under which it fails to do so.