发表机构
Nanjing University; Ant Group; HSBC Business School, Peking University(南京大学; 蚂蚁集团; 北京大学汇丰商学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出模糊感知多智能体框架AMAO,通过两阶段对齐处理运筹问题描述的逻辑不一致性,提升自动建模的准确率与鲁棒性。
AI 中文摘要
在现实场景中,运筹学(OR)问题通常由利益相关者以不完整知识和模糊表达来描述。此类模糊描述无法直接用于建模。因此,自动化运筹问题求解需要事先处理这种逻辑不一致性。为应对这一挑战,我们提出了一个用于自动运筹问题求解的模糊感知多智能体框架AMAO,该框架首先通过两阶段对齐处理逻辑不一致性,然后再进行下游建模和编码。具体而言,一个具有OR引导专家结构的逻辑对齐智能体首先通过变量、参数、目标和约束上的监督路由生成对齐候选。随后,一个源充分性验证器判断原始描述是否支持所需的修复。当证据不足时,一个交互式修复智能体请求额外信息并在建模和编码前修订描述。此外,提出了一个模糊感知基准及其匹配的对话扩展以支持两个阶段的评估。32B模型实现了86.8%的错误恢复成功率和60.4%的求解准确率,分别达到基础模型分数的1.98倍和1.86倍。8B模型实现了47.9%的求解准确率,超过了最佳评估变体如DeepSeek-V4-Flash和Claude-sonnet-4.5。交互式修复在需要澄清的情况下进一步实现了90.71%的完整描述准确率。这些结果支持将基于上下文的修复与证据评估和针对性用户澄清相结合用于自动化OR建模。
英文摘要
Operations research (OR) problems are often described by stakeholders with incomplete knowledge and vague expressions in real-world settings. Such ambiguous descriptions can not be used to formulation directly. Thus, automating OR problems solving requires processing this logical inconsistency in advance. To address this challenge, we propose an \textbf{A}mbiguity-Aware \textbf{M}ulti-Agent Framework for \textbf{A}utomated \textbf{O}R Problem Solving, \textbf{AMAO}, which first addresses logical inconsistency through two-stage alignment before downstream formulation and coding. Specifically, a logical alignment agent with OR-guided experts structure first produces an aligned candidate using supervised routing over variables, parameters, objectives, and constraints. Then, a source-sufficiency verifier determines whether the original description supports the required repairs. When evidence is insufficient, an interactive repair agent requests additional information and revises the description before modeling and coding. In addition, an ambiguity-aware benchmark and its matched dialogue extension are proposed to support evaluation of both stages. The 32B model achieved 86.8\% error-recovery success and 60.4\% solution accuracy, reaching 1.98 and 1.86 times the respective base-model scores. The 8B model achieved 47.9\% solution accuracy, exceeding the best evaluated variant such as DeepSeek-V4-Flash and Claude-sonnet-4.5. Interactive repair further achieved 90.71\% complete-description accuracy when cases requiring clarification. These results support combining context-based repair with evidence assessment and targeted user clarification for automated OR modeling.