发表机构
Cardinal Operations; Shanghai Jiao Tong University(杉数科技; 上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出OR-Clarify基准与InterOPT两阶段框架,解决LLM建模前的问题描述不完整问题,提升槽位恢复性能,将OR辅助转化为选择性完整性决策。
AI 中文摘要
大型语言模型(LLM)越来越多地被用于将自然语言问题描述转化为优化模型,但现实中的运筹学(OR)请求往往不完整:缺失目标、约束或业务规则会改变最终的数学规划模型。现有评估大多假设问题描述完整,因此忽略了智能体在建模前是否知晓需要澄清的情况。我们推出OR-Clarify,这是一个用于预建模澄清的基准。每个任务提供部分公开问题描述,隐藏结构化的槽位,并通过与模拟用户的有限交互来评估智能体。该基准支持开放式和选择式两种澄清方式,测量槽位恢复、停止行为、隐性假设和交互成本。我们还提出Interactive Optimization(InterOPT),这是一个两阶段框架,可识别未解决的建模关键缺口,并据此指导是否提出下一个问题或停止。在选择式实验中,InterOPT在精确槽位恢复方面显著优于所有基线方法;在开放式设置中,它仍与强大的现有方法具有竞争力。总体而言,OR-Clarify和InterOPT将OR辅助重新定义为选择性完整性决策:需要时进行澄清,准备好时停止,并量化仍缺失的内容。
英文摘要
Large language models (LLMs) are increasingly used to formulate optimization models from natural-language problem descriptions, yet realistic operations research (OR) requests are often incomplete: missing objectives, constraints, or business rules can change the resulting mathematical program. Existing evaluations largely assume a complete specification and therefore overlook whether an agent knows when clarification is needed before modeling. We introduce OR-Clarify, a benchmark for pre-formulation clarification. Each task presents a partial public problem description, withholds structured hidden slots, and evaluates agents through bounded interaction with a simulated user. The benchmark supports both openended and choice-based clarification, and measures slot recovery, stopping behavior, silent assumptions, and interaction cost. We further propose Interactive Optimization (InterOPT), a two-stage framework that identifies unresolved formulation-critical gaps and uses them to guide whether to ask the next question or to stop. In our choice-based experiments, InterOPT substantially outperforms all baselines in exact slot recovery; in the open-ended setting, it remains competitive with strong prior methods. Together, OR-Clarify and InterOPT reframe OR assistance as a selective completeness decision: clarify when needed, stop when ready, and quantify what remains missing.
Comments16 pages, 4 figures, 4 tables. Revised exposition and added references; results unchanged