输入凸神经网络作为数学优化中的代理模型
Input convex neural networks as surrogates in mathematical optimisation
浏览论文内容
中文总结 AI 辅助
该研究提出用输入凸神经网络(ICNN)作为数学优化的代理模型,其能得到更紧致的线性规划松弛,开发了分支定界算法,在三类案例中验证了其精度与FNN相当且求解更高效。
中文摘要 AI 辅助
将训练好的神经网络作为代理嵌入优化问题是运筹学中的成熟做法。主流方法采用带ReLU激活函数的前馈神经网络(FNN),其分段线性结构可实现精确的混合整数规划(MIP)重构,但随着网络规模增大,计算量会变得庞大。当底层响应近似为凸或凹函数时,我们提倡使用输入凸神经网络(ICNN)作为结构更优的代理模型。这种凸架构具备两项计算优势:其一,ICNN的MIP重构形式往往比FNN的MIP重构能得到更紧致的线性规划(LP)松弛,在有利情况下不存在整数间隙;其二,ICNN可通过ReLU激活函数的上境图表示实现基于LP的重构,尽管该嵌入并非总是精确的。当嵌入不精确时,我们利用ICNN的特性在箱型定义域上构造最强的连续松弛,即ICNN图像的凸包,其下界为上境图,上界为凹包;该构造在输入凸性条件下是易处理的,但对于一般ReLU网络则较难实现。在此基础上,我们开发了一种分支定界算法,该算法在每个节点构建此松弛,直接对输入变量进行分支,而非像MIP重构那样对中间变量分支,且当上境图嵌入有效时可在根节点终止。针对人道主义粮食援助、油井路径规划和葡萄酒调配的案例研究表明,ICNN代理模型的精度与FNN相当,且在求解时间和可扩展性上有所提升,这支持当底层函数为凸、凹或可近似为凸/凹时,将ICNN作为默认代理模型。
英文摘要
Embedding trained neural networks as surrogates within optimisation problems is an established practice in operations research. The prevailing approach uses feedforward neural networks (FNNs) with ReLU activations, whose piecewise-linear structure admits an exact but computationally intensive mixed-integer programming (MIP) reformulation as the networks grow. We advocate input convex neural networks (ICNNs) as structurally superior surrogates when the underlying response is approximately convex or concave. The convex architecture offers two computational advantages. First, the ICNN-MIP formulation tends to yield a tighter linear programming (LP) relaxation than its FNN-MIP counterpart, with no integrality gap in favourable instances. Second, ICNNs uniquely admit an LP-based reformulation via epigraph representations of ReLU activations, though this embedding is not always exact. When it is not, we exploit the properties of ICNNs to construct the strongest continuous relaxation over box domains, namely, the convex hull of the ICNN's graph, bounded below by the epigraph and above by the concave envelope; this construction is tractable under input convexity but hard for general ReLU networks. On this basis, we develop a branch-and-bound algorithm that builds this relaxation at each node, branches directly on input variables rather than intermediate variables as in MIP reformulations, and terminates at the root node whenever the epigraph embedding is valid. Case studies on humanitarian food aid, oil well routing, and wine blending show that ICNN surrogates match FNN accuracy and deliver gains in solve time and scalability, supporting ICNN as the default surrogate when the underlying function is convex, concave, or well-approximated as such.
发表机构
- Aalto University(阿尔托大学)
- KTH Royal Institute of Technology(瑞典皇家理工学院)
- Technical University of Denmark(丹麦技术大学)
机构由 AI 辅助整理,请以论文原文为准。