发表机构
Ironproof(Ironproof)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
IRONPROOF是基于SMT等价性检查的COBOL转Python工具,在2345个COBOL文件上验证等价性,独立程序等价率52.6%,可检测溢出,证明仅保证中间表示与Python等价而非COBOL实现等价。
AI 中文摘要
将COBOL程序翻译为现代编程语言可能会改变程序的计算结果,而大型语言模型(LLM)生成的翻译无法保证等价性。本文提出IRONPROOF工具,该工具先将COBOL解析为中间表示,再生成Python代码,将两者编码为基于共享输入的Z3公式,最终输出机器可验证的等价性证书(UNSAT)或反例(SAT)。在2345个COBOL文件(含GnuCOBOL测试、NIST CCVS85、开源集合及本文生成或编写的程序)中,782个进入检查流程:606个(77.5%)被证明等价,101个部分验证,无被反驳的程序,75个在流程中失败未计入统计。在独立编写的程序上,等价率为52.6%(291个中的153个),而在本文编写的程序上等价率达92.3%。在153个独立证明中,37个覆盖编码器未建模输入的程序单次执行;所有606个证明中仅14个对已证明输出所依赖的输入进行量化。在两个公开业务应用语料库(AWS CardDemo和IBM GenApp)上,编码器未对任何程序进行端到端建模。2026年4月仅用LLM的基准测试在94个独立编写的程序中54个失败,统计的是编码器无法验证的翻译。PIC边界溢出检测在其可分析的程序中标记出22.1%的程序存在超出范围的输出,此为上限。此前与GnuCOBOL 3.2.0的对比发现,49个可执行独立证明中有24个在证明域内的变量与运行时结果不一致。证明仅能确定生成的Python计算结果与中间表示所表示的COBOL计算结果一致,而非与COBOL实现等价。
英文摘要
Translating COBOL to a modern language can change what a program computes, and LLM translations carry no guarantee of equivalence. We present IRONPROOF, which parses COBOL into an intermediate representation, generates Python, encodes both as Z3 formulas over shared inputs, and emits either a machine-checkable equivalence certificate (UNSAT) or a counterexample (SAT). On 2,345 COBOL files (GnuCOBOL tests, NIST CCVS85, open-source collections, and programs we generated or wrote), 782 enter the checking path: 606 (77.5%) are proved equivalent, 101 are partially verified, none is refuted, and 75 fail inside our pipeline and stay in the denominator. On independently authored programs the rate is 52.6% (153 of 291), against 92.3% on programs we wrote. Of the 153 independent proofs, 37 cover a single execution of a program that reads input the encoder does not model, and across all 606 proofs only 14 quantify over an input that a proved output depends on. On two public business-application corpora (AWS CardDemo and IBM GenApp), the encoder models no program end to end. An April 2026 LLM-only baseline failed on 54 of 94 independently authored programs, counting translations our encoder could not verify. PIC-bounded overflow detection flags an out-of-range output in 22.1% of the programs it can analyze, an upper bound. An earlier confrontation with GnuCOBOL 3.2.0 found 24 of 49 executable independent proofs disagreeing with the runtime on a variable in the proof's domain. A proof establishes that the generated Python computes what our intermediate representation says the COBOL computes, not equivalence to a COBOL implementation.
Comments13 pages, 6 tables, 1 figure. Code and data: https://github.com/dom-omg/ironproof-cobol