发表机构
College of Computing and Data Science, Nanyang Technological University; School of Computing and Information Systems, Singapore Management University; Tsinghua University(南洋理工大学计算与数据科学学院; 新加坡管理大学计算与信息系统学院; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FormuEvo是LLM引导的进化框架,将MIP模型设计转化为符号空间的进化优化,结合求解器感知诊断与结构化内存,发现的MIP模型性能远超专家设计及现有LLM方法,求解加速最高达5.5倍且知识可跨场景迁移。
AI 中文摘要
混合整数规划(MIP)是运筹学与工业优化的核心。尽管大型语言模型(LLM)近期在从自然语言自动构建MIP模型方面展现出潜力,但它们更侧重语义正确性,却忽视了模型的强度,严重制约了下游求解器的效率。我们提出FormuEvo,一种由LLM引导的进化框架,用于自动发现求解器高效的MIP模型。FormuEvo将MIP模型设计视为对MIP模型符号空间的进化优化,该空间以可执行建模程序表示,通过LLM驱动的交叉、变异与修复操作迭代生成、评估并选择更优候选方案。为超越盲目探索,FormuEvo引入了求解器感知的诊断机制,利用细粒度的求解器统计数据作为语言梯度以实现针对性优化。此外,结构化内存将过往经验抽象为可复用的建模策略,避免冗余探索,同时实现对未知问题的零样本迁移并为较小规模的LLM提供引导。针对各类线性与非线性问题的实验表明,FormuEvo发现的模型显著优于专家设计的模型及现有基于LLM的方法,可将求解器速度提升高达5.5倍,且提炼的知识可在不同问题与模型规模间有效迁移。
英文摘要
Mixed-integer programming (MIP) lies at the core of operations research and industrial optimization. While large language models (LLMs) have recently shown promise in automated MIP modeling from natural language, they prioritize semantic correctness but overlook formulation strength, severely bottlenecking the efficiency of downstream solvers. We propose FormuEvo, an LLM-guided evolutionary framework for automated discovery of solver-efficient MIP formulations. FormuEvo frames MIP formulation design as evolutionary optimization over the symbolic space of MIP formulations, represented as executable modeling programs, by iteratively generating, evaluating, and selecting stronger candidates via LLM-driven crossover, mutation, and repair operations. To move beyond blind exploration, FormuEvo introduces a solver-informed diagnosis mechanism that exploits fine-grained solver statistics as verbal gradients for targeted refinement. Additionally, a structured memory abstracts prior experience into reusable modeling strategies, avoiding redundant exploration while enabling zero-shot transfer to unseen problems and bootstrapping smaller LLMs. Experiments across diverse linear and non-linear problems demonstrate that FormuEvo discovers formulations that significantly outperform both expert-designed formulations and existing LLM-based approaches, accelerating solvers by up to 5.5$\times$, with distilled knowledge transferring effectively across problems and model scales.
Comments27 pages, 6 figures, and 9 tables. To appear in the Proceedings of EMNLP 2026