发表机构
EPAM Systems(EPAM系统公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出形式化语义块模型与基于执行判断的基准,验证其在Oracle-PostgreSQL迁移规范中的有效性,支持确定性为形式化概念但非LLM实现者的独立经验质量指标。
AI 中文摘要
本工作提出了一种用于规范的形式化语义块模型,以及一种与模型能力无关的、用于评估规范质量的基于执行判断的基准。规范被表示为包含语义块、依赖关系、块所属规则、决策点和显式开放问题的结构,需满足四项机器可检查的良构性条件:无环性、单一所有权、约束支配性以及总体性或歧义停止。确定性在模型理论中被定义为所有符合规范的实现之间的一致性,并通过独立实现者之间的收敛性进行经验估计。该模型被实例化到一个包含18个块和19条依赖边的Oracle到PostgreSQL迁移规范中。计算验证显示,五层分解通过依赖闭包将平均每任务上下文减少了约71%,覆盖了研究定义的Oracle构造分类的85.5%,所有识别出的缺口均已分类,且不受测试的替代划分的帕累托支配,从用于定义原始结构的引文衍生边中以99.9百分位数被恢复。基准保持实现者面板固定,包含一个强制性的无规范对照组,并使用PostgreSQL 16和一个实时Oracle实例作为确定性执行判断。六项设计研究,包括三项预注册操作和三项诊断分析,进一步检验规范的影响。对25个单元子样本的重复运行揭示了一个经验变异性下限,其中中位臂差值范围为14.4个百分点。结果支持确定性作为一种形式化概念,但不支持其作为评估的当代LLM实现者的独立经验质量指标。
英文摘要
This work introduces a formal semantic-block model for specifications and an execution-judged benchmark for evaluating specification quality independently of model capability. A specification is represented as a structure comprising semantic blocks, dependency relations, block-owned rules, decision points, and explicitly open questions, subject to four machine-checkable well-formedness conditions: acyclicity, single ownership, constraint domination, and totality or ambiguity-stop. Determinacy is defined model-theoretically as agreement among all conforming implementations and is estimated empirically through convergence across independent implementers. The model is instantiated on an Oracle-to-PostgreSQL migration specification containing 18 blocks and 19 dependency edges. Computational validation shows that the five-layer decomposition reduces mean per-task context by approximately 71% through dependency closures, covers 85.5% of the study-defined Oracle construct taxonomy with all identified gaps triaged, is not Pareto-dominated by the tested alternative partitions, and is recovered at the 99.9th percentile from citation-derived edges not used to define the original structure. The benchmark keeps the implementer panel fixed, includes a mandatory no-specification control arm, and uses PostgreSQL 16 and a live Oracle instance as deterministic execution judges. Six designed studies, including three pre-registered manipulations and three diagnostic analyses, further examine specification effects. Repeated runs on a 25-unit subsample reveal an empirical variability floor with a median arm-delta spread of 14.4 percentage points. The results support determinacy as a formal concept but not as a standalone empirical quality metric for the evaluated contemporary LLM implementers.
Comments11 pages, 1 figure, 7 tables, 37 references