发表机构
School of Computer Science and Technology, University of the Chinese Academy of Sciences(中国科学院大学计算机科学与技术学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
IR2Solve以结构化中间表示为核心构建优化自动建模流水线,仅用1次LLM语义调用,在保证目标函数正确性的同时大幅降低推理成本,性能优于Chain-of-Experts和SAC-Opt等系统。
AI 中文摘要
大型语言模型(LLMs)可将自然语言优化问题转换为求解器可用的建模形式,但直接生成代码存在缺陷:模式、索引和语义错误会导致编译失败、模型不可行或目标函数不正确,而迭代修复、搜索和多智能体工作流会增加推理成本。本文提出IR2Solve,一种以中间表示为核心的自动建模流水线,它仅通过一次语义LLM调用生成受模式约束的ModelIR,随后执行两个确定性阶段:验证和IR到求解器的编译。ModelIR使用受限的类Python表达式字符串明确表示集合、参数、变量、目标函数和约束。具体的标量约束约定将有限的按索引约束族表示为单独条目,减少自由索引和隐式量化错误,同时简化下游验证和编译。在6个经过清洗的优化基准测试中,IR2Solve实现了出色的目标函数正确性,且与近期的优化建模系统具有竞争力。对153个IndustryOR和ComplexLP实例的受控消融实验显示,结构化IR接口、标量约束指令和确定性验证带来了连续的性能提升。在匹配的10个实例成本面板上,IR2Solve每个实例仅使用1次语义调用,而Chain-of-Experts和SAC-Opt每个实例分别使用8次和39次调用,消耗的token量分别为IR2Solve的3.3倍和22.9倍。这些结果表明,结构化中间表示结合确定性的后生成处理,为基于LLM的优化自动建模提供了实用的精度-成本权衡方案。
英文摘要
Large language models (LLMs) can translate natural-language optimization problems into solver-ready formulations, but direct code generation is brittle: schema, indexing, and semantic errors can cause compilation failures, infeasible models, or incorrect objectives, while iterative repair, search, and multi-agent workflows increase inference cost. We present IR2Solve, an intermediate-representation-first autoformulation pipeline that uses a single semantic LLM call to produce a schema-constrained ModelIR, followed by two deterministic stages: verification and IR-to-solver compilation. ModelIR explicitly represents sets, parameters, variables, objectives, and constraints using restricted Python-like expression strings. A concrete scalar-constraint convention represents finite per-index constraint families as individual entries, reducing free-index and implicit-quantification errors while simplifying downstream verification and compilation. Across six cleaned optimization benchmarks, IR2Solve achieves strong objective correctness and remains competitive with recent optimization-modeling systems. A controlled ablation on 153 IndustryOR and ComplexLP instances shows sequential gains from the structured IR interface, the scalar-constraint instruction, and deterministic verification. On a matched ten-instance cost panel, IR2Solve uses one semantic call per instance, whereas Chain-of-Experts and SAC-Opt use 8 and 39 calls per instance and consume 3.3 and 22.9 times the token volume of IR2Solve, respectively. These results show that structured intermediate representations, combined with deterministic post-generation processing, provide a practical accuracy-cost trade-off for LLM-based optimization autoformulation.
Comments17 pages, 3 figures, 13 tables