发表机构
Nanjing University(南京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对高维贝叶斯优化中现有LLM与智能体方法失效的问题,提出HERA智能体,通过假设与证据引导搜索并自适应配置策略,在合成与真实任务上优于对比方法。
AI 中文摘要
高维贝叶斯优化(HDBO)旨在当变量数量相对于评估预算较大时实现样本高效的优化。近期基于大语言模型(LLM)和智能体的贝叶斯优化方法在运行过程中融入任务知识并调整搜索决策,但主要已在低维和中维问题上进行了评估。我们探究这一范式能否迁移到更高维的领域。我们的实验表明,这些方法在高维领域无法保持可靠性,在该领域中,挑战不仅在于在何处进行评估,还在于当目标函数的有用结构未知时,应采用何种建模假设和搜索几何。因此,我们引入了HERA,一种基于假设与证据引导的研究智能体,它利用任务上下文、优化反馈和结构诊断来修订搜索假设、选择并配置HDBO策略,并确定其执行长度。其数值优化引擎PRISM在每个搜索块内顺序生成并评估候选点,并在每次观测后更新数值模型。HERA在四个无元数据的合成函数上,与强大的数值HDBO基线保持竞争力,并优于所评估的基于LLM和智能体的方法。在八个真实世界任务中,HERA在大多数基准上取得了所有评估系统中最优的平均最终目标值。进一步分析表明,结构诊断会改变策略使用,元数据效应因任务而异,而自适应搜索块降低了推理成本。
英文摘要
High-dimensional Bayesian optimization (HDBO) seeks sample-efficient optimization when the number of variables is large relative to the evaluation budget. Recent LLM-based and agentic BO methods incorporate task knowledge and adapt search decisions during a run, but have primarily been evaluated on low- and moderate-dimensional problems. We ask whether this paradigm can transfer to the higher-dimensional regime. Our experiments show that these methods do not remain reliable in the high-dimensional regime, where the challenge is not only where to evaluate, but also which modeling assumption and search geometry to use when the objective's useful structure is unknown. We therefore introduce HERA, a Hypothesis- and Evidence-guided Research Agent that uses task context, optimization feedback, and structural diagnostics to revise search hypotheses, select and configure HDBO strategies, and determine their execution length. PRISM, its numerical optimization engine, generates and evaluates candidates sequentially within each search block, updating numerical models after each observation. HERA remains competitive with strong numerical HDBO baselines and outperforms the evaluated LLM-based and agentic methods on four metadata-free synthetic functions. Across eight real-world tasks, HERA achieves the best mean final objective among all evaluated systems on most benchmarks. Further analyses show that structural diagnostics change strategy use, metadata effects vary across tasks, and adaptive search blocks reduce inference cost.