arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向运筹学的基于模拟的不确定性感知大语言模型推理

Uncertainty-Aware Simulation-Based Inference for Operations Research with Large Language Models

Liang Guo, Lin Shaochong, Shen Zuo-Jun Max, Zhang Kun

arXiv 2608.00019首次发表:更新:

发表机构

Institute of Statistics and Big Data, Renmin University of China; The University of Hong Kong; School of Information, Renmin University of China(中国人民大学统计与大数据研究院; 香港大学; 中国人民大学信息学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对大语言模型用于运筹学建模时的短视策略易引发下游错误的问题,提出一种不确定性感知无训练推理框架,通过短前瞻模拟与重要性重采样提升OR公式生成的可靠性,在多基准上表现优于基线。

AI 中文摘要

将大语言模型(LLM)应用于运筹学(OR)任务仍具挑战性,因为其正确性依赖连贯的建模过程,而非仅正确的最终答案。标准自回归生成采用短视策略,有时无法预判部分公式能否有效扩展为全局一致的优化模型,导致局部合理步骤可能引发下游公式或求解器代码的灾难性错误。为解决该问题,本文提出一种用于OR数学建模的不确定性感知、无训练推理框架。无需更新模型参数,该方法通过短前瞻模拟评估中间候选步骤,量化下游预测不确定性或概率集中度;随后通过重要性重采样动态选择更可能生成连贯数学公式的候选。在多个OR基准(包括NL4OPT、MAMO及IndustryOR)上的实证评估表明,该框架始终优于标准基线和低温基线,为可靠的OR公式生成建立了高效的无训练范式。

英文摘要

Deploying large language models (LLMs) for operations research (OR) tasks remains challenging because correctness depends on a coherent modeling process, not merely a correct final answer. Standard autoregressive generation operates on a myopic policy, which sometimes fails to anticipate whether a partial formulation can be validly extended into a globally consistent optimization model. Consequently, locally plausible steps may propagate into catastrophic downstream formulation or solver code errors. To address this, we propose an uncertainty-aware, training-free inference framework for OR mathematical modeling. Without updating model parameters, our method evaluates intermediate candidate steps using short lookahead simulations to quantify downstream predictive uncertainty or probability concentration. Candidates that demonstrate a higher likelihood of yielding coherent mathematical formulations are then dynamically selected via importance resampling. Empirical evaluations across multiple OR benchmarks (including NL4OPT, MAMO, and IndustryOR) demonstrate that our framework consistently outperforms both standard and low-temperature baselines, establishing an efficient, training-free paradigm for reliable OR formulation generation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑