发表机构
State Key Laboratory of Novel Software Technology, Nanjing University; School of Artificial Intelligence, Nanjing University(南京大学现代软件工程国家重点实验室; 南京大学人工智能学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究推出跨领域智能体BBO基准AgenticBBO-Bench,实验显示智能体BBO性能优于直接LLM方法及多数数值优化器,还分析了影响性能的因素并评估了多款LLM在前沿挑战中的表现。
AI 中文摘要
黑盒优化(BBO)出现在许多目标评估成本高昂且有限的科学与工程问题中。近期的大语言模型(LLM)智能体通过结合任务语义、计算资源、优化工具以及反馈驱动的决策制定,为解决BBO提供了新途径,由于与数学严谨工具的集成,展现出巨大潜力。然而,现有的智能体BBO研究使用不同的任务领域和系统配置,导致其结果难以比较,且难以分离出单个设计选择的影响。因此,我们推出AgenticBBO-Bench,这是一个针对智能体BBO的跨领域基准,涵盖合成函数、超参数优化、数据库调优、芯片设计和分子设计,采用统一的有限预算评估协议。在我们的实验中,智能体BBO在所有五个领域的家族平均得分均高于直接基于LLM的方法,并且在四个领域中优于最佳数值优化器。我们进一步研究了影响智能体性能的三个因素:优化工具、任务信息与先验知识,以及LLM在搜索过程中的作用。我们的结果表明,额外的数值工具并不能持续提升性能,任务语义具有广泛的实用性,而更具体的先验知识可靠性较低,数值优化器能够有效吸收智能体建立的搜索轨迹带来的增益。最后,我们在AgenticBBO-Bench中推出了一个包含五项任务的前沿挑战,并在Codex智能体框架下评估了七个LLM,其中GPT-6 Astra和DeepSeek-V4.1-Flash在评估模型的性能与成本的帕累托前沿上。我们的代码可在该https URL获取。
英文摘要
Black-box optimization (BBO) arises in many scientific and engineering problems where objective evaluations are expensive and limited. Recent large language model (LLM) agents offer a new way to approach BBO by combining task semantics, computation, optimization tools, and feedback-driven decision making, showing great potential due to the integration with mathematically rigorous tools. However, existing agentic BBO studies use different task domains and system configurations, making their results difficult to compare and the effects of individual design choices hard to isolate. We therefore introduce AgenticBBO-Bench, a cross-domain benchmark for agentic BBO spanning synthetic functions, hyperparameter optimization, database tuning, chip design, and molecular design under a unified finite-budget evaluation protocol. In our experiments, agentic BBO achieves higher family-averaged scores than direct LLM-based methods in all five domains and outperforms the best numerical optimizers in four. We further study three factors shaping agent performance: optimization tools, task information and prior knowledge, and the role of the LLM during search. Our results show that additional numerical tools do not consistently improve performance, task semantics are broadly useful while more specific priors are less reliable, and numerical optimizers can effectively absorb gains from search trajectories established by the agent. Finally, we introduce a five-task frontier challenge within AgenticBBO-Bench and evaluate seven LLMs under the Codex agent harness, where GPT-6 Astra and DeepSeek-V4.1-Flash lie on the Pareto frontier of performance and cost among the evaluated models. Our code is available at https://github.com/lamda-bbo/agentic-bbo.