发表机构
Harvard University(哈佛大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出D-IMPL,一种基于扩散模型的参数化黑箱优化求解器,通过学习最小化策略来分摊计算成本,并具有样本复杂度保证和广泛实证验证。
AI 中文摘要
扩散模型在多个领域的生成建模任务中展现出强大的能力,表现出从样本中学习复杂分布的显著能力。在本文中,我们利用这种能力来设计一种高效的、通用的基于扩散模型的参数化黑箱优化(BBO)求解器,其中优化器在学习阶段只能对目标函数的查询进行黑箱访问,但能够在推理阶段减少每个BBO实例的额外计算成本,同时捕捉非凸目标的潜在多模态景观。为了将我们的公式转化为兼容的生成建模任务,我们引入了最小化策略的概念作为一种新的解决方案概念,它定义了在解空间上的采样分布,该分布应集中在每个BBO实例的最小化器集合周围。然后,我们提出了基于扩散的迭代最小化策略学习(D-IMPL),一种实用的基于生成模型的求解器,用于解决参数化BBO,它采用扩散模型来学习最小化策略,其密度与目标值的负指数成正比,从而在不同BBO实例之间分摊计算成本。此外,我们通过建立样本复杂度保证来展示D-IMPL算法的性能,该保证表明可以在$O(\log(1/\delta))$次迭代内有效学习到$\delta$-近似最小化策略,并通过在一系列有约束和无约束BBO任务上的广泛实证评估来验证。
英文摘要
Diffusion models have demonstrated strong power in generative modeling tasks across multiple domains, exhibiting a remarkable capability of learning complex distributions from samples. In this paper, we leverage such capability to design an efficient universal diffusion-based solver for parameterized black-box optimizations (BBO), where the optimizer has only black-box access to queries of the objective function at the learning stage, yet is able to reduce the additional computational cost at the inference stage for each BBO instance while also capturing the potential multi-modal landscape of non-convex objectives. To cast our formulation as a compatible generative modeling task, we introduce the notion of minimization policy as a new solution concept, which defines a sampling distribution over the solutions that should concentrate around the minimizer set for each BBO instance. We then propose Diffusion-based Iterative Minimization Policy Learning (D-IMPL), a practical generative-model-based solver for solving parameterized BBOs that employs diffusion models to learn a minimization policy, whose density is proportional to the exponential of the negated objective values, thereby amortizing the computational costs across different BBO instances. Furthermore, we demonstrate the performance of our D-IMPL algorithm by establishing a sample complexity guarantee showing that a $δ$-approximate minimization policy can be effectively learned within $O(\log(1/δ))$ iterations, and by extensive empirical evaluations over a range of constrained and unconstrained BBO tasks.
Comments12 pages, 4 figures, 3 tables