AI 中文总结
研究旨在解决优化建模问题,提出PEARL系统,通过在循环中使用Python执行和数学编程求解器,学习优化策略,在多轮工具集成设置下提升求解率,在优化建模任务中表现优于强大基线。
AI 中文摘要
优化建模是将通常用自然语言描述的现实世界决策问题转化为形式化数学公式和可执行求解器代码的过程。虽然大语言模型的进展有望自动化此过程,但现有方法大多是一次性的。我们引入了PEARL,一个用于交互式优化建模的系统,它在循环中使用Python执行和数学编程求解器。PEARL学习何时测试部分模型、如何根据求解器诊断进行修正以及何时停止。在多轮工具集成设置中运行,利用中间执行结果等改进公式和求解器代码。在各种优化基准测试中,PEARL显著提高了验证求解率,优于强一次性和工具增强基线。
英文摘要
Optimization modeling is the process of translating real-world decision problems, often described in natural language, into formal mathematical formulations and executable solver code. While recent advances in large language models have shown promise in automating this process, most existing approaches remain one-shot: a model produces a formulation once, without executing it, conditioning on solver feedback, or iteratively revising errors. This stands in sharp contrast to real-world optimization modeling, which is inherently interactive and proceeds through repeated solve-debug-revise cycles. We introduce PEARL, a system for interactive optimization modeling that uses Python execution and mathematical programming solvers inside this loop. Rather than relying on a fixed repair workflow, PEARL learns when to test partial models, how to revise from solver diagnostics, and when to stop. It operates in a multi-turn tool-integrated setting where intermediate execution results, feasibility signals, and solution checks are used to improve both formulations and solver code before finalization. Across diverse optimization benchmarks, PEARL substantially improves verified solve rates over strong one-shot and tool-augmented baselines; notably, our PEARL-Qwen3-\textbf{4B} model outperforms the much larger DeepSeek-V3.2-\textbf{685B} in both macro- and micro-averaged accuracy on optimization modeling tasks.