arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SOVER:通过大语言模型辅助的SMT验证实现优化重述的形式化认证

SOVER: Formal Certification of Optimization Reformulations via LLM-Assisted SMT Verification

Swapnil Bhattacharyya, Mayank Baranwal

arXiv 2609.00728首次发表:更新:

发表机构

TCS Research; Indian Institute of Technology Bombay(塔塔咨询服务研究中心; 印度理工学院孟买分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出LLM辅助的SMT框架SOVER,结合Z3与dReal工具,在NLEquiv-150基准上以99.33%准确率验证优化问题重述的等价性,解决经验验证不可靠的问题。

AI 中文摘要

大语言模型(LLM)在跨建模语言翻译和重述复杂数学优化问题方面展现出显著潜力,但仅通过经验求解器执行验证此类变换不可靠,因为求解器结果可能受局部最小值、结构性超时、数值伪影以及不同表述间微妙语义分歧的影响。我们提出SOVER,这是一种LLM辅助的SMT框架,将语义映射与形式化认证分离:Z3检查混合整数线性表述的域交叉可行性和全局目标序保持性,dReal为连续非线性表述提供容差感知的可行性/范围和ε-argmin检查。我们还提出NLEquiv-150,这是一个包含100个等价和50个故意构造的困难非等价非线性重述对的公开基准。借助LLM提取的映射,SOVER对150对中的149对(准确率99.33%)进行了正确分类,包括全部50个困难负例;唯一错误源于不完整的映射提取。

英文摘要

Large Language Models (LLMs) have shown remarkable promise in translating and reformulating complex mathematical optimization problems across modeling languages. However, validating such transformations through empirical solver executions alone is unreliable, as solver outcomes may be affected by local minima, structural timeouts, numerical artifacts, and subtle semantic divergence between formulations. We introduce SOVER, an LLM-assisted SMT framework that separates semantic mapping from formal certification: Z3 checks domain cross-feasibility and global objective-order preservation for mixed-integer linear formulations, while dReal provides tolerance-aware feasibility/range and $ε$-argmin checks for continuous nonlinear formulations. We also introduce NLEquiv-150, a public benchmark of 100 equivalent and 50 deliberately hard non-equivalent nonlinear reformulation pairs. With LLM-extracted mappings, SOVER classifies 149/150 pairs (99.33%) correctly, including all 50 hard negatives; the sole error is an incomplete mapping extraction.

CommentsAccepted to EMNLP 2026 Findings

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑