arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

区域缩减ReLU神经网络的优化

Optimization with Region-Reduced ReLU Neural Networks

Christoph Plate, Caroline Ganzer, Mirko Hahn, Alexander Klimek, Heyuan Liu, Sebastian Sager, Kai Sundmacher, Hanna Wilhelm

arXiv 2609.31380首次发表:更新:

发表机构

Max Planck Institute for Dynamics of Complex Technical Systems; Otto von Guericke University; Leibniz University Hannover(马克斯·普朗克复杂技术系统动力学研究所; 奥托·冯·格里克大学; 汉诺威莱布尼茨大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出区域缩减模型压缩方法,通过稳定不稳定神经元和合并冗余神经元减少线性区域,在保持预测精度下,使超结构优化计算时间降低40%-50%,解误差小于1%。

AI 中文摘要

涉及整数决策和ReLU激活神经网络(ReLU ANN)的数学模型优化是一项具有挑战性的任务。然而,此类模型在许多应用领域中是一种使能技术。一个突出的例子是化学工程中的超结构优化,其中ReLU ANN经常被用作复杂非线性过程的替代模型。我们调查了这一领域的最新进展。我们认为,除了ANN的网络规模和训练选项外,ReLU激活几何形状以及感兴趣域上的线性区域数量对计算优化性能有强烈影响。虽然标准模型压缩方法(如结构化剪枝)减少了网络规模,但它们并未明确解决几何考虑。因此,我们提出了一种新颖的“区域缩减”模型压缩方法,该方法结合了不稳定神经元的稳定化和冗余神经元的合并,以减少线性区域的数量,同时通过误差补偿保持预测准确性。我们在多个优化用例上评估了我们的方法与标准压缩方法的对比。首先,是二维峰值函数,我们可以可视化其激活几何形状。其次,是对三个化学过程的单个替代ReLU ANN的优化,第三,是涉及三个ReLU ANN、额外过程子模型和二元变量的混合超结构优化问题。超结构问题的结果表明,区域缩减具有巨大潜力,计算时间减少了40%至50%,在获得压缩模型的努力可忽略的情况下,所得解与参考解的差距小于1%。

英文摘要

Optimization of mathematical models involving integer decisions and neural networks with ReLU activation (ReLU ANNs) is a challenging task. Nevertheless, such models are an enabling technology in many application domains. A prominent example is superstructure optimization in chemical engineering, where ReLU ANNs are frequently employed as surrogate models for complex nonlinear processes. We survey recent developments in this area. We argue that in addition to network size and training options of the ANNs, the ReLU activation geometry and the number of linear regions on the domain of interest have a strong impact on computational optimization performance. While standard model compression approaches such as structured pruning reduce network size, they do not explicitly address geometric considerations. Therefore, we propose a novel \emph{region-reduced} model compression approach that combines the stabilization of unstable neurons and the merging of redundant neurons to reduce the number of linear regions while maintaining predictive accuracy through error compensation. We evaluate our method against standard compression approaches on multiple optimization use cases. First, the two-dimensional peaks function for which we can visualize the activation geometry. Second, on optimization over individual surrogate ReLU ANNs for three chemical processes, and third, on a hybrid superstructure optimization problem that involves the three ReLU ANNs, additional process submodels, and binary variables. The results for the superstructure problem demonstrate the large potential of region reduction with a decrease of 40\% to 50\% in computational time, yielding solutions closer than 1\% to the reference at negligible effort of obtaining the compressed model.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑