arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从死代码和静态需求到可运行的引擎:基于编码智能体的软件复兴

From Dead Code and Static Requirements to Working Engines: Software Revival with Coding Agents

Tianyu Liu, Dingyuan Dai, Yufan Du, Zhen Yang

arXiv 2609.36161首次发表:更新:

发表机构

Tsinghua University; UCLA(清华大学; 加州大学洛杉矶分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出ReviveBench基准,包含复兴与重建两类任务,通过隐藏验证器评估编码智能体恢复软件和重建引擎的能力,并揭示验证器缺陷,提供实用检查方法。

AI 中文摘要

编码智能体能否在保留软件底层方法的同时,恢复不再运行的软件,并根据开放规范重建工业软件引擎?为此,我们引入了ReviveBench,一个包含两个任务族类的基准测试,通过针对原生执行环境、既有工程工具或专门构建的参考实现进行校准的隐藏验证器进行评估。复兴任务族包括十个涉及依赖不兼容、核心模块删除、遗留构建以及基于GPU的基础模型的任务。每个起始工作区均无法通过验证,而评估的最强模型在至少一次运行中通过了全部十个任务。在污染控制实验中,标识符混淆将行级相似度从0.51--0.96降至0.03--0.44,且未降低任何评估模型的通过率。在知识截止日期之后创建的代码库上,最强模型在九次运行中通过了八次。重建任务族包含十三个任务,涵盖数值、几何、硬件和事务系统(如CAD和CRM)。两个模型在全部十三个任务上达到了基准的通过标准,尽管我们的审计表明,CFD任务无法建立数值求解器能力。基准构建和审计发现了28个验证器缺陷,包括24个假阴性和2个假阳性。这些发现表明,可执行验证本身可能引入显著的测量误差。我们提出了三项实用检查:测试规定方法能否达到评分阈值,调查独立生成候选之间的一致性,以及根据提交的工件重新计算诊断指标。因此,ReviveBench既提供了对软件复兴和引擎重建的评估,也提供了验证用于测量编码智能体在软件设计方面能力的验证器的案例。

英文摘要

Can coding agents restore software that no longer runs while preserving its underlying methods, and reconstruct industrial software engines from open specifications? Here we introduce ReviveBench, a benchmark with two task families evaluated by hidden verifiers calibrated against native execution environments, established engineering tools, or purpose-built reference implementations. The revival family comprises ten tasks involving dependency incompatibilities, deleted core modules, legacy builds, and a GPU-based foundation model. Every starting workspace fails verification, and the strongest evaluated model passes all ten tasks in at least one run each. In contamination-control experiments, identifier obfuscation reduces line similarity to the original implementations from 0.51--0.96 to 0.03--0.44 without reducing the observed pass rate of any evaluated model. On repositories created after the stated knowledge cutoffs, the strongest model passes eight of nine runs. The reconstruction family comprises thirteen tasks spanning numerical, geometric, hardware, and transactional systems (e.g. CAD and CRM). Two models meet the benchmark's pass criteria on all thirteen, although our audit shows that the CFD task cannot establish numerical-solver capability. Benchmark construction and auditing uncover 28 verifier defects, including 24 false negatives and two false positives. These findings show that executable verification can itself introduce substantial measurement error. We present three practical checks: test whether prescribed methods can reach the grading thresholds, investigate agreement among independently generated candidates, and recompute diagnostics from submitted artifacts. ReviveBench thus provides both an evaluation of software revival and engine reconstruction, and cases in validating the verifiers used to measure coding agents for software design.

Comments14 pages, 3 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑