arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BBOWP-Bench:评估大型语言模型在黑盒优化文字问题上的表现

BBOWP-Bench: Evaluating LLMs on Black-Box Optimization Word Problems

Yutaro Yamada, Kei Hiroshima, Nozomu Yoshinari, Kento Uchida, Shinichi Shirakawa

arXiv 2608.02612首次发表:更新:

发表机构

Yokohama National University(横滨国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出BBOWP问题设定并构建BBOWP-Bench基准,评估发现现有LLMs可根据评估预算选合适算法,但搜索空间设计仍存不足。

AI 中文摘要

优化问题的建模方式会强烈影响最终解的质量,但良好的建模通常需要深厚的专业知识。因此,近期研究探索了如何从自然语言描述中自动推导优化问题,不过现有基准测试聚焦于目标和约束可明确写为数学表达式的场景。许多具有实际重要性的问题天然属于黑盒优化(Black-Box Optimization,BBO)问题,这类问题仅可观测目标值,函数形式不可用。在BBO中,作为问题建模一部分的搜索空间设计,以及优化算法的选择,对问题求解至关重要。用大型语言模型(Large Language Models,LLMs)自动化这些过程是一项重大挑战。本文提出了黑盒优化文字问题(Black-Box Optimization Word Problems,BBOWP)这一新的问题设定,要求系统从黑盒优化任务的自然语言描述中同时推断搜索空间和优化算法。为支持该设定的研究,我们构建了BBOWP基准套件(BBOWP-Bench),这是一个针对BBOWP的数据集和评估框架。每个实例包含自然语言问题描述、可执行评估环境以及人工设计的基准建模,可同时评估搜索空间设计和算法选择。利用该基准,我们首次对LLMs进行了评估,结果显示当前LLMs能够根据给定的评估预算选择合适的算法,但有时在搜索空间设计上存在困难,尤其是在问题描述信息不足或搜索空间具有高度问题特异性时,难以识别重要变量并平衡其范围。我们的代码和数据集可在此httpsURL获取。

英文摘要

Formulating an optimization problem strongly affects the quality of the final solution, yet good formulations usually require substantial expertise. Recent studies have therefore examined how to automatically derive optimization problems from natural-language descriptions, but existing benchmarks focus on settings where objectives and constraints can be written explicitly as mathematical expressions. Many practically important problems are naturally treated as black-box optimization (BBO) problems, in which only objective values are observable, and the functional form is unavailable. In BBO, the search space design, a part of the problem formulation, and the selection of the optimization algorithm are crucial for problem-solving. Automating these processes with large language models (LLMs) is a significant challenge. This paper introduces Black-Box Optimization Word Problems (BBOWP), a novel problem setting in which a system must infer both a search space and an optimization algorithm from a natural-language description of a black-box optimization task. To support research on this setting, we establish the BBOWP Benchmark Suite (BBOWP-Bench), a dataset and evaluation framework for BBOWP. Each instance combines a natural-language problem description, an executable evaluation environment, and a human-designed baseline formulation, allowing evaluation of both search-space design and algorithm selection. Using this benchmark, we provide the first evaluation of LLMs and show that current LLMs are capable of selecting suitable algorithms based on the given evaluation budget. However, they sometimes struggle with search space design, particularly in identifying important variables and balancing their ranges when the problem description is less informative or the search space is highly problem-specific. Our code and dataset are available at https://github.com/shiralab/bbowp-bench.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑