发表机构
Scientific Computing Center (SCC), Karlsruhe Institute of Technology (KIT)(卡尔斯鲁厄理工学院科学计算中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出IQS-BO,一种基于先验数据拟合网络的上下文内查询选择方法,通过监督学习在合成先验上学习查询决策,以低成本实现贝叶斯优化,并匹配或超越现有方法。
AI 中文摘要
贝叶斯优化(BO)是优化昂贵黑盒函数的强大框架,但通常需要在每次评估步骤重新拟合代理模型并最大化采集函数。基于先验数据拟合网络(PFNs)的上下文内方法通过在合成先验上预训练变换器来分摊部分成本。PFNs4BO分摊了代理模型,但仍依赖于数值最大化的采集函数,而FIBO通过从学习密度中采样优化器位置来完全在上下文内执行BO,这固定了决策规则且不采用代理模型。学习采集函数使用训练好的网络对有限候选集进行评分,但由于缺乏查询标签,通过强化学习在先前解决的任务上学习评分。我们提出IQS-BO,一种通过在合成先验上进行监督学习来学习查询决策的PFN。在单次前向传播中,IQS-BO预测每个候选在集合上最大化目标的概率,并且我们证明其目标的最小化器是该事件的后验概率。该模型可以在没有代理模型的情况下预训练以进行完全上下文内BO,或者将固定概率代理模型的预测作为额外输入,仅分摊决策步骤。我们的方法以采集方法成本的一小部分提出查询,同时在合成和真实世界基准上匹配或超越使用高斯过程(GPs)的标准BO以及可用的上下文内方法。最后,我们提出了一种用于预训练PFNs的混合先验,该先验结合了来自GPs的样本与具有扭曲输入、孤立窄最优或平台区的函数,这些函数难以被GP代理中常见的平稳核建模。我们表明,在此先验上预训练可以带来改进的优化性能。
英文摘要
Bayesian Optimisation (BO) is a powerful framework for the optimisation of expensive black-box functions, but typically requires refitting a surrogate and maximising an acquisition function at every evaluation step. In-context approaches based on Prior-data Fitted Networks (PFNs) amortise part of this cost by pre-training transformers on functions drawn from synthetic priors. PFNs4BO amortises the surrogate but still relies on a numerically maximised acquisition function, while FIBO performs BO fully in-context by sampling optimiser locations from a learned density, which fixes the decision rule and admits no surrogate. Learned acquisition functions score a finite candidate set with a trained network, but, lacking a label for the query, learn the score by reinforcement learning on previously solved tasks. We propose IQS-BO, a PFN that learns the query decision by supervised learning on synthetic priors. In a single forward pass, IQS-BO predicts the probability that each candidate maximises the objective over the set, and we show that the minimiser of its objective is the posterior probability of this event. The model can be pre-trained without a surrogate for fully in-context BO, or take the predictions of a fixed probabilistic surrogate as additional input, amortising only the decision step. Our method proposes queries at a fraction of the cost of acquisition-based methods, while either matching or outperforming standard BO with Gaussian processes (GPs) and available in-context methods on synthetic and real-world benchmarks. Finally, we propose a mixture prior for pre-training PFNs which combines samples from GPs with functions exhibiting warped inputs, isolated narrow optima, or plateaus that are poorly modeled by stationary kernels common in GP surrogates. We show that pre-training on this prior can lead to improved optimisation performance.