arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ModelEquivBench:LLM生成优化模型的多关系验证评估系统

ModelEquivBench: Certifying Multi-Relational Evaluation of LLM-Generated Optimization Models

Penglin Zhu, Jungang Xu

arXiv 2607.29431首次发表:更新:

AI 中文总结

该研究提出多关系评估系统ModelEquivBench,将其用于评估GPT-5.4等三个LLM生成的优化模型,发现不同模型在特征不同阶段失败,无法简化为单一准确率分数。

AI 中文摘要

大型语言模型(LLM)越来越多地从自然语言生成优化模型,但现有评估通常将生成的模型及其真值简化为单一的等价/不等价判定或执行成功率——这些标签既无法独立验证,也无法忠实反映两种公式达成一致的多种不同情形。我们提出ModelEquivBench,这是一种可验证的多关系评估系统,会报告每对的语义特征E0至E6:模型构建与精确导入(E0)、经验证的表示对齐(E1)、同空间及投影可行集关系(E2、E3)、目标序等价(E4)、最优值相等(E5)、优化器集等价(E6)。每个已判定条目都带有与关系适配、可独立复核的证据:E0至E1对应可复现的轨迹或显式映射,E2至E6的肯定结论对应精确有理证书,否定结论对应显式见证。不完整的映射搜索、不受支持的结构及资源限制会产生类型化的UNKNOWN或N/A结果,而非猜测;未满足的前提则报告为ABSENT。使用ModelEquivBench在同一组173个基础问题(每个模型对应346个单元)、无修复协议的条件下评估三个模型快照——GPT-5.4、Claude Sonnet 4.6和Qwen3.5-397B-A17B,所得特征揭示了粗粒度基线无法体现的差异:49、35和25个单元包含可执行候选,但在至少一个支持的关系上被判定为否定;在E2验证了经映射的可行集相等的对中,分别出现25、8和18次结构拒绝。这三个模型快照在特征的不同阶段失败,因此无法有意义地简化为单一准确率分数。

英文摘要

Large language models increasingly generate optimization models from natural language, but existing evaluation often reduces a generated model and its ground truth to a single equivalent/not-equivalent verdict or an execution-success rate--labels that are neither independently checkable nor faithful to the multiple distinct senses in which two formulations can agree. We present ModelEquivBench, a certifying, multi-relational evaluation system that reports a per-pair semantic profile E0--E6: model construction and exact ingestion (E0), verified representation alignment (E1), same-space and projected feasible-set relations (E2, E3), objective-order equivalence (E4), optimal-value equality (E5), and optimizer-set equivalence (E6). Each decided entry carries relation-appropriate, independently re-checkable evidence: replayable traces or explicit maps for E0--E1, exact-rational certificates for positive E2--E6 conclusions, and explicit witnesses for supported negatives. Incomplete mapping search, unsupported structure, and resource limits produce typed UNKNOWN or N/A outcomes rather than guesses, while unmet prerequisites are reported as ABSENT. Using ModelEquivBench to evaluate three model snapshots--GPT-5.4, Claude Sonnet 4.6, and Qwen3.5-397B-A17B--on the same frozen cohort of 173 base problems (346 cells per model) under a no-repair protocol, the resulting profiles expose distinctions that coarse baselines do not represent: 49, 35, and 25 cells contain executable candidates that are nevertheless certified negative on at least one supported relation, and 25, 8, and 18 structural rejections occur on pairs for which E2 certifies mapped feasible-set equality under a verified map. The three model snapshots fail at different stages of the profile and therefore cannot be meaningfully reduced to a single accuracy score.

Comments9 pages, 2 figures, 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑