arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.11180cs.AIcs.SE

SemVerBench:基准测试LLM对版本约束解析语义的理解

SemVerBench: Benchmarking LLM Comprehension of Version-Constraint Resolution Semantics

  • Brown University(布朗大学)

机构由 AI 辅助整理,请以论文原文为准。

Qibai Chen, Zeming Liu

AI总结:

提出SemVerBench基准,评估LLM在npm、PEP 440和Cargo上的版本约束解析,发现系统性盲点,建议编码代理委托解析器而非自行推理。

AI中文摘要:

大型语言模型(LLM)编码代理经常需要判断某个版本是否满足诸如^1.2.3或>=2.0,<3之类的约束,然而它们对版本约束语义的掌握从未被直接测量过。我们引入了SemVerBench,这是首个针对三个生态系统(npm、PEP 440、Cargo)的LLM版本约束解析语义基准测试:包含240个具有唯一答案、可由机器检查的项目,这些项目以作者中立的方式从四个均衡来源(每个生态系统的官方测试套件加上三个前沿LLM提议者)构建,并由一个非循环的双实现预言机进行标注。评估六个前沿模型后,我们发现了系统性、可预测的按机制盲点:部分比较器进位规则(>1.2意味着>=1.3.0)使每个模型在Cargo上陷入困境(接近60%),尽管标准的PEP 440前缀匹配是普遍的,但在零填充/发布后边界情况下,GPT-5.1崩溃(0/26),而Claude保持在97-100%(在67项预言机验证集上验证)。Opus显著优于所有其他模型,Sonnet优于OpenAI模型(McNemar检验)。这些失败看起来更像是激活/应用差距而非知识差距:注入规则或一个轻微的正确提示可以恢复大多数错误,而区间分解则不能,并且模型在同一规则的基本形式上达到上限。作者分层分析发现没有统计上显著的自我偏袒。由于该任务是可验证的,并且存在一个免费、100%正确的解析器,工具委托达到约100%:编码代理应将版本解析委托给解析器,而不是在头脑中推理版本。

英文摘要:

Large language model (LLM) coding agents constantly decide whether a version satisfies a constraint such as ^1.2.3 or >=2.0,<3, yet their grasp of version-constraint semantics has never been measured directly. We introduce SemVerBench, the first benchmark of LLM version-constraint resolution semantics across three ecosystems (npm, PEP 440, Cargo): 240 machine-checkable items with unique answers, built author-neutrally from four balanced sources (each ecosystem's official test suite plus three frontier LLM proposers) and labeled by a non-circular two-implementation oracle. Evaluating six frontier models, we find systematic, predictable per-mechanism blind spots: a partial-comparator carry rule (>1.2 means >=1.3.0) traps every model on Cargo (near 60%), and although standard PEP 440 prefix matching is universal, on zero-pad/post-release corner cases GPT-5.1 collapses (0/26) while Claude stays at 97-100% (verified on a 67-item oracle-validated set). Opus significantly outperforms all other models, and Sonnet outperforms the OpenAI models (McNemar). The failures look more like an activation/application gap than a knowledge gap: injecting the rule or a light correct hint recovers most errors, whereas interval decomposition does not, and models are at ceiling on the basic forms of the same rules. An author-stratified analysis finds no statistically significant self-favoritism. Because the task is verifiable and a free, 100%-correct resolver exists, tool delegation reaches ~100%: coding agents should delegate version resolution to a resolver rather than reason about versions in-head.

补充信息

↑