正确答案,昂贵模型:基于LLM的优化建模中的效率差距
Right Answers, Costly Models: The Efficiency Gap in LLM-based Optimization Modeling
浏览论文内容
中文总结 AI 辅助
本研究提出OptDachshund框架和EfficientOpt基准,系统评估LLM在优化建模中的效率,发现大多数正确模型求解时间更长,强调需同时评估正确性与计算效率。
中文摘要 AI 辅助
优化建模将现实世界中的决策问题表述为数学规划,求解器可利用这些规划来找到最优决策。大型语言模型(LLMs)可以自动化这一过程,但由此产生的正确表述在构建和求解时可能需要大量的时间和内存,从而限制了实际的可扩展性。因此,我们系统地研究了LLMs能否从自然语言描述中识别问题结构,并应用合适的优化建模技术来生成能够正确且高效地求解问题的数学模型和求解器代码。为此,我们首先整理了OptTips,这是一个包含八个家族中50种专家建模技术的知识库。利用这一知识,我们开发了OptDachshund,这是一个多智能体框架,可将现有优化基准中的问题转化为新任务,用于评估LLMs对建模技术的运用。它为同一任务和数据构建了常规模型和专家数学模型及求解器代码,为正确性和计算成本提供了基线。由此产生的EfficientOpt基准包含561个经专家审查的任务,并配有配对的参考实现。对11个代表性LLMs的评估揭示了一个在可比较测量下正确求解任务上的效率差距:对于每个LLM,大多数生成的程序比专家对应程序需要更长的求解时间。在可比较的参考规模子集内,57%的具有正确目标值且变量和线性约束更少的程序记录了更长的求解器时间。案例研究表明,不同的建模技术可以在相似的记录成本下实现相同的最优值。如果代码在准备数据和构建模型上花费更长时间,更快的求解可能不会减少执行时间。因此,LLM优化建模应同时评估正确性和计算效率。
英文摘要
Optimization modeling formulates real-world decision problems as mathematical programs that solvers can use to find optimal decisions. Large language models (LLMs) can automate this process, but the resulting correct formulations can require substantial time and memory to construct and solve, limiting practical scalability. Therefore, we systematically investigate whether LLMs can identify problem structure from natural-language descriptions and apply suitable optimization modeling techniques to generate mathematical models and solver code that solve the problems correctly and efficiently. To this end, we first curate OptTips, a knowledge base of 50 expert modeling techniques in eight families. Using this knowledge, we develop OptDachshund, a multi-agent framework that transforms problems from existing optimization benchmarks into new tasks for evaluating LLMs' use of modeling techniques. It constructs conventional and expert mathematical models with solver code for the same task and data, providing baselines for correctness and computational cost. The resulting EfficientOpt benchmark contains 561 expert-reviewed tasks with paired reference implementations. Evaluation of 11 representative LLMs reveals an efficiency gap on correctly solved tasks with comparable measurements: for every LLM, most generated programs take longer to solve than their expert counterparts. Within the comparable reference-size subset, 57\% of programs with correct objective values and fewer variables and linear constraints have longer recorded solver times. Case studies show that different modeling techniques can achieve the same optimal value at similar recorded cost. Faster solving may not reduce execution time if the code takes longer to prepare data and build the model. LLM optimization modeling should therefore be evaluated for both correctness and computational efficiency.
发表机构
- Great Bay University(大湾区大学)
- Beijing Jiaotong University(北京交通大学)
- Beihang University(北京航空航天大学)
- Peking University(北京大学)
机构由 AI 辅助整理,请以论文原文为准。