arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于证伪的大语言模型生成优化模型验证:可靠的测试组及其检测极限

Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits

Haifeng Li, Mo Hai

arXiv 2607.16646首次发表:更新:

发表机构

School of Information, Central University of Finance and Economics(信息学院,中央财经大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究大语言模型生成优化模型的验证问题,基于证伪理论开发测试组,通过求解器调用测试候选模型,实验证实该测试组可靠,能检测多种错误,误报率低,再现可检测性模式。

AI 中文摘要

大语言模型如今能将决策问题的自然语言描述转换为求解器就绪的优化模型,但可能会悄然出错,生成的模型运行时也可能会制定错误的问题。本文针对此情况开发了一种基于证伪的验证理论。描述中的每个数值量都是一个类型化插槽,候选模型仅通过对插槽转换实例的求解器调用来测试,不参考参考模型或标签。从对偶性、比较静态分析和多面体极限论证中,我们得出了一组测试类,涵盖方向、曲率、挤压探针、禁止极限、湮灭和交换。每个测试都是可靠的,因此违反测试可证明不忠实,误报率设计为零。我们描述了这种验证永远无法看到的情况,给出了确定检测到规范错误类的条件,并证明没有固定阈值扰动测试器能同时可靠且非平凡。对来自NL4OPT的326个真实模型和四个基准系列的实验证实了该理论。该测试组的误报率为0.0%,而阈值测试器为54.9%,检测到70.0%的已认证条件类突变体,判定40.4%执行精度评分不可见的突变体有罪,并再现了预测的可检测性模式,包括其零值。

英文摘要

Large language models now translate natural-language descriptions of decision problems into solver-ready optimization models, and they fail silently. A generated model often runs and still encodes the wrong problem, while standard evaluation compares optimal values against labeled answers that deployment does not provide. How to certify such a model without any reference is the question this paper addresses. We develop falsification-based verification. Every numeric quantity in a problem description plays a role that the text itself states, such as a capacity, a requirement, or a unit cost, and any correct model must respond to changes in these quantities as the stated roles dictate. From duality and sensitivity analysis we derive a battery of solver-based tests that are individually sound, so a violation certifies a faulty model and the false-positive rate is zero by design. We characterize the errors that no test of this kind can see, give conditions under which each canonical error class is detected with certainty, and prove that perturbation testers with tuned thresholds cannot be simultaneously sound and nontrivial. Across 326 ground-truth models, a synthetic family, and four public benchmarks with two generators, the battery flags 0.0% of faithful models while a threshold tester flags 54.9%; it detects 56.1% of core formulation errors, 70.0% under certified preconditions, and 40.4% of the errors that value-based scoring provably cannot see, and it reproduces the predicted detectability pattern including its blind spots. Every flag carries a machine-checkable certificate that localizes the defect, and a full audit costs about 25 millisecond-scale solver calls per model. Classical sensitivity analysis and duality thus offer a rigorous, label-free audit that complements existing evaluation of AI-generated optimization models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑