发表机构
Beihang University; The Chinese University of Hong Kong, Shenzhen; Cardinal Operations Technology Co.(北京航空航天大学; 香港中文大学(深圳); 卡达克科技有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SemOPT通过语义奖励模型和分层奖励引导搜索,修复LLM优化建模中的语义错误,在七个基准上平均准确率提升7.6%。
AI 中文摘要
运筹学支持能源、经济、医疗等领域的决策制定。解决运筹学问题通常始于优化建模,即将自然语言问题描述转化为可执行的求解器代码。LLM为自动化这一过程提供了有前景的途径,但仍容易出错。在实践中,这些错误可分为两类:语法错误指求解器代码无法成功运行或被求解器判定为不可行;语义错误指求解器代码成功返回目标值但违背了原始问题的意图。由于语义错误不会触发运行时故障,因此难以检测和纠正。为解决这一问题,我们提出了SemOPT,一个用于纠正基于LLM的优化模型的语义引导框架。SemOPT结合了一个语义奖励模型,用于区分忠实的数学模型与看似合理但错误的模型,以及一个自适应修正系统,该系统在建模空间上应用分层奖励引导搜索。在七个优化建模基准上的实验表明,SemOPT达到了新的最先进水平,并在复杂数据集上相较于最强基线实现了平均7.6%的准确率提升。
英文摘要
Operations research supports decision-making in domains such as energy, economics, and healthcare. Solving operations research problems typically begins with optimization modeling, which translates a natural-language problem description into executable solver code. LLMs offer a promising way to automate this process, but they remain prone to errors. In practice, these errors can be divided into two categories: syntactic errors refer to solver code that fails to run successfully or is judged infeasible by the solver; semantic errors refer to solver code that successfully returns an objective value but violates the intent of the original problem. Since semantic errors do not trigger runtime failures, they are difficult to detect and rectify. To address this problem, we introduce SemOPT, a semantic-guided framework for correcting LLM-based optimization models. SemOPT combines a semantic reward model that distinguishes faithful math models from plausible but incorrect ones with an adaptive correction system that applies hierarchical reward-guided search over the modeling space. Experiments on seven optimization modeling benchmarks show that SemOPT establishes a new state of the art and achieves an average 7.6% accuracy improvement over the strongest baseline on complex datasets.
CommentsAccepted at EMNLP 2026 (Findings)