AI 中文总结
该研究针对优化器系统(SOSO)的模拟优化,提出利用内部优化几何结构的 PRIME 求解器,在多类测试平台上实现了优异性能,大幅降低了方差,为模拟优化提供了新方向。
AI 中文摘要
我们研究优化器系统(SOSO)的模拟优化:这是一种基于智能体的模拟,其中每个智能体在每个决策 epoch 都要解决一个结构化优化问题——线性规划(LP)、混合整数规划或动态规划。此类系统出现在供应链、电力市场和物流领域,但标准模拟优化将模拟视为黑箱,忽略了内部优化的几何结构。我们对 SOSO 进行形式化,并开发了一个框架,将这种几何结构转化为计算优势。我们证明,内部 LP 的最优基和对偶变量会在动态过程中传播,从而得到外部目标的精确、无偏、单副本微小扰动分析(IPA)梯度,仅存在测度为零的基变化作为阻碍。我们推导了由正向模拟中可计算的基不一致计数控制的共同随机数协方差界。我们进一步证明,IPA 方差随反馈深度呈指数增长,形式化了牛鞭效应,并引入了带误差预算的决策替代,将 LP 时域和强化学习替代统一在一个界下。我们将这些整合为 PRIME,这是一种随机近似求解器,集成了 IPA 梯度、自适应步长和多起始空间多样化。在覆盖 SOSO 分类的六个测试平台上,PRIME 在相同预算下达到最佳或并列最佳的最优性差距,且振荡接近零、种子间方差小;IPA 相比独立有限差分实现了约 2000 倍的每副本方差降低,而共同随机数在 1000 个 SKU、六个配送中心的供应链中,将配对差异方差降低了超过 10000 倍。这些结果表明,利用嵌入优化的几何结构是模拟优化中具有实际意义的重要方向。
英文摘要
We study simulation-optimization of systems of optimizers (SOSO): agent-based simulations in which every agent solves a structured optimization - a linear program (LP), mixed-integer program, or dynamic program - at every decision epoch. Such systems arise in supply chains, electricity markets, and logistics, yet standard simulation optimization treats the simulation as a black box, discarding the inner optimization's geometry. We formalize SOSO and develop a framework that converts this geometry into computational advantage. We prove that the inner LP's optimal basis and dual variables propagate through the dynamics to yield an exact, unbiased, single-replication infinitesimal perturbation analysis (IPA) gradient of the outer objective, with measure-zero basis changes as the only obstructions. We derive a common-random-numbers covariance bound governed by a computable basis-disagreement count from the forward simulation. We further prove that IPA variance grows exponentially with feedback depth, formalizing the bullwhip effect, and introduce surrogate-as-decision with an error budget unifying LP-horizon and reinforcement-learning surrogates under one bound. We compose these into PRIME, a stochastic-approximation solver integrating the IPA gradient, adaptive step sizing, and multi-start spatial diversification. On six testbeds spanning the SOSO taxonomy, PRIME achieves the best or tied-best optimality gap at equal budget with near-zero oscillation and narrow seed-to-seed variance; IPA yields a ~2000x per-replication variance reduction over independent finite differences, and common random numbers cut paired-difference variance by over 10,000x in a 1,000-SKU, six-distribution-center supply chain. These results establish exploiting embedded-optimization geometry as a practically significant direction for simulation optimization.