DASH:用于自动化贝叶斯优化的解耦自适应代理-采集调控器
DASH: Decoupled Adaptive Surrogate - Acquisition Harness for Automated Bayesian Optimization
AI总结:
DASH是用于LLM增强AutoBO的解耦自适应调控器,通过分离代理与采集的适配逻辑,在四项化学优化任务中较最佳AutoBO基线实现了12.51%的轨迹级加速因子提升和5.00%的端点增强因子提升。
AI中文摘要:
贝叶斯优化(BO)依赖于代理模型和采集函数,但最适合的选择会因任务和优化阶段而异。自动化贝叶斯优化(AutoBO)通过在线适配BO组件解决这种变异性。然而,现有的AutoBO方法要么仅适配一个组件,导致另一个组件不匹配并形成瓶颈,要么在统一准则下联合选择代理-采集对,忽略了它们的不同作用:代理选择取决于预测可靠性,而采集适配应响应优化进程。本文提出DASH,即用于大语言模型(LLM)增强型AutoBO的解耦自适应代理-采集调控器。DASH通过预测可靠性、不确定性校准和排序一致性选择代理;其两阶段采集控制器定期在采集函数间重新分配配额,据此构建BO候选列表,并将最终选择委托给LLM。DASH还包含集成调控器,由知识引导的热启动和结构化记忆组成,以将优化建立在领域知识和累积反馈的基础上。在四项化学优化任务中,DASH在轨迹级加速因子上优于最佳AutoBO基线12.51%,在端点增强因子上优于5.00%。在不同LLM主干上结果均表现强劲, ablation( ablation 即消融实验,保留缩写)验证了所有组件的互补贡献。全表和行为污染检查未发现可检测证据,表明直接基准记忆或源单元泄漏可解释这些增益。
英文摘要:
Bayesian optimization (BO) relies on a surrogate model and an acquisition function, yet the most suitable choices vary across tasks and optimization stages. Automated Bayesian optimization (AutoBO) addresses this variability by adapting BO components online. However, existing AutoBO methods either adapt one component, leaving the other mismatched and creating a bottleneck, or jointly select surrogate--acquisition pairs under a shared criterion, overlooking their distinct roles: surrogate selection depends on predictive reliability, whereas acquisition adaptation should respond to campaign context.In this paper, we propose DASH, a Decoupled Adaptive Surrogate--Acquisition Harness for large-language- model (LLM)-enhanced AutoBO. DASH selects surrogates by predictive reliability, uncertainty calibration, and ranking consistency; its two-stage acquisition controller periodically reallocates quotas across acquisition functions, builds a BO shortlist accordingly, and delegates final selection to an LLM. DASH also incorporates an integrated harness, consisting of knowledge-guided warm start and structured memory, to ground optimization in domain knowledge and accumulated feedback. Across four chemical optimization tasks, DASH outperforms the best AutoBO baseline by 12.51% in trajectory-level Acceleration Factor and 5.00% in endpoint Enhancement Factor. Results remain strong across LLM backbones, and ablations verify the complementary contributions of all components. Full-table and behavioral contamination checks find no detectable evidence that direct benchmark memorization or source-cell leakage explains these gains.