LLM 的终点与可靠决策的起点
Where the LLM Ends and Reliable Decisions Begin
- LogiModel AI
- UT Austin(德克萨斯大学奥斯汀分校)
- MIT(麻省理工学院)
- UC Berkeley(加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对优化问题自然语言转代码中语言模型过度使用的问题,提出 ANVIL 架构,仅用一次语言模型生成 LaTeX,再由确定性编译器翻译,在 NLP4LP 基准上编译成功率达 92.9%,准确率高达 98.9%(简单)和 91.1%(困难)。
AI中文摘要:
将优化问题的自然语言描述转化为求解器可执行代码的系统,通常在每个阶段都使用语言模型,包括从数学公式到可执行建模代码的最终翻译。我们提出了 ANVIL 编译器架构,在该架构中我们分离了这些关注点。语言模型仅被调用一次,并辅以基于问题类型的约束引导,以生成 LaTeX 公式。随后,一个确定性编译器将该 LaTeX 翻译为代码,全程不涉及语言模型。我们描述了该确定性编译器(一个归一化器、一个生成类型化中间表示的递归下降解析器、将符号绑定到数据集模式的分析遍及一个代码生成器),并在 NLP4LP 基准的 354 个简单和困难问题上对其进行了评估。该编译器为 354 个公式中的 329 个(92.9%)生成了代码,在所有这些情况下都采用了确定性路径,从未回退到模型生成的代码。中位编译时间低于我们计时器的 10 毫秒分辨率。总体而言,我们的公式在简单问题上达到了 98.9% 的准确率,在困难问题上达到了 91.1% 的准确率。编译成功率和准确率之间的差距是分析的关键点,我们对此进行了分析:未能编译的公式、返回不可行的问题、运行时出错的问题以及返回错误目标的问题。几乎所有这些错误都追溯到公式本身,而非翻译过程。ANVIL 通过纯粹在语言模型有效的领域使用它们,而非将其作为万能工具,在领先的基准上表现出色。
英文摘要:
Systems that turn natural-language descriptions of optimization problems into solver-ready code generally use a language model at every stage, including the final translation from a mathematical formulation into executable model-building code. We propose the ANVIL compiler architecture, where we separate these concerns. A language model is called only once, assisted by constraint guidance based on problem type, to produce a LaTeX formulation. A deterministic compiler then translates that LaTeX into code with no language model involvement. We describe the deterministic compiler (a normalizer, a recursive-descent parser producing a typed intermediate representation, analysis passes that bind symbols to a dataset schema, and a code emitter) and evaluate it on the 354 easy and hard problems of the NLP4LP benchmark. The compiler produced code for 329 of 354 formulations (92.9%), taking the deterministic path in every one of those cases and never falling back to model-generated code. Median compile time was below the 10ms resolution of our timer. Overall, our formulations achieved an accuracy of 98.9% over easy problems and 91.1% for hard problems. The gap between these compilation and accuracy figures is a key point of analysis, and we analyze it: formulations that failed to compile, problems that returned as infeasible, problems raising errors at runtime, and problems returning a wrong objective. Almost all of these errors trace back to the formulation rather than to the translation. ANVIL performs exceptionally well on leading benchmarks by using language models purely where they are effective, rather than as a catch-all tool.