发表机构
Stanford University; Johns Hopkins University; Massachusetts Institute of Technology(斯坦福大学; 约翰斯·霍普金斯大学; 麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出 DCEmbed 方法,利用神经网络的差凸表示避免二进制变量,通过迭代凸凹过程高效求解嵌入神经代理的优化问题,在多个实验中显著加速收敛并提升解质量。
AI 中文摘要
神经代理模型通过用高效的 learned 近似替代昂贵或难以处理的模型组件,可以加速大规模优化,但解决由此产生的嵌入问题可能仍然代价高昂。例如,具有 ReLU 激活函数的神经网络的标准精确编码允许通过混合整数求解器解决问题,但会添加大量二进制变量以适应激活函数的非线性,这可能使问题在计算上变得不可行。为了解决这个问题,我们提出了 DCEmbed,一种针对嵌入神经代理的优化问题的启发式方法,它利用网络的差凸(DC)表示并避免添加激活二进制变量。利用 ReLU 网络的 DC 表示中的共享结构,我们推导出其凸分量的缩减规模精确公式,该公式可以仅使用两个线性不等式和每个隐藏神经元一个连续辅助变量嵌入到优化问题中。使用此方法,通过迭代惩罚凸凹过程求解问题,其中每一步仅逼近神经项的凹部分。原始目标、约束和任何离散决策都被保留,允许标准凸或混合整数优化求解器在每次迭代中联合优化宿主和代理。在二次规划、混合整数资源分配和神经两阶段随机规划的实验中,我们的方法在向高质量可行解的进展方面比使用精确混合整数嵌入的方法快得多。特别是,DCEmbed 在资源分配问题上实现了比最佳精确基线低 4 倍的归一化原始积分,而在两阶段随机规划中,它比 Gurobi ML 快约 5 倍达到全局代理最优。
英文摘要
Neural surrogates can accelerate large-scale optimization by replacing expensive or intractable model components with efficient learned approximations, but solving the resulting embedded problems can remain prohibitively costly. For instance, standard exact encodings of neural networks with ReLU activations allow the problem to be solved by mixed-integer solvers, but add large numbers of binary variables to accommodate the nonlinearity of the activations, which can render the problem computationally prohibitive. To address this, we propose DCEmbed, a heuristic for optimization problems with embedded neural surrogates that leverages the difference-of-convex (DC) representation of the network and avoids adding activation binaries. Exploiting shared structure within the DC representation of a ReLU network, we derive a reduced-size, exact formulation for its convex components that can be embedded in optimization problems using just two linear inequalities and one continuous auxiliary variable per hidden neuron. Using this, the problem is solved via an iterative penalty convex-concave procedure, where only the concave portions of the neural terms are approximated at each stage. The original objective, constraints, and any discrete decisions are retained, allowing standard convex or mixed-integer optimization solvers to optimize the host and surrogate jointly at each iteration. In experiments on quadratic programs, mixed-integer resource allocation, and neural two-stage stochastic programming, our method demonstrates much faster progress toward high-quality feasible solutions than approaches using exact mixed-integer embeddings. In particular, DCEmbed achieves $4\times$ lower normalized primal integral than the best exact baseline on the resource allocation problem, while in two-stage stochastic programming it reaches the global surrogate optimum $\sim 5\times$ faster than Gurobi ML.