识别代码的隐式声明式表示以辅助代码仓库迁移
Identifying Latent Declarative Representations of Code for Assisting Repository Migration
浏览论文内容
中文总结 AI 辅助
该研究针对遗留代码仓库迁移难题,提出 ADFD-Migrate 方法,通过显式化代码的隐式声明式表示提升移植效果,在新基准 f2x50 上取得优于对比方法的移植正确性与完整性。
中文摘要 AI 辅助
遗留软件仓库在未加注释的代码中嵌入了数十年的领域知识,这使得代码的理解与现代化改造变得困难。我们将程序视为其计算过程的隐式声明式描述的实现,探究是否将这种隐式声明式表示显式化可提升仓库规模的代码移植效果。ADFD-Migrate 用带注释的数据流图(ADFD)近似表示该隐式表示,其中涵盖了进程、数据存储、外部实体、数据流及行为契约。大型语言模型(LLM)在静态分析覆盖检查的引导下,从受限的仓库上下文推断源 ADFD。依赖感知分块对受限的进程组排序,以生成目标语言代码。随后,源 ADFD 与静态恢复的目标 ADFD 之间的差异会指导代码的重新生成。我们在 f2x50 上评估 ADFD-Migrate,f2x50 是一个包含 50 个 Fortran 仓库的新基准,代码行数介于 1500 至 160 万行之间,涵盖三个复杂度层级。我们从两个维度评估生成的移植结果:移植正确性,通过源预言机的行为一致性衡量;移植完整性,通过综合迁移结果指数衡量。在 382 个精心挑选的 Fortran 预言机探测项中,生成的 Python 代码通过了 327 个(占比 85.6%),其中 40 个仓库通过了所有尝试的探测项。ADFD-Migrate 可将全部 382 个计划行为暴露为可运行目标,而直接翻译和仓库上下文翻译分别仅能暴露 99 和 98 个,静态轮廓和依赖分块的 ablation 方法则分别仅能暴露 69 和 30 个。它还实现了 93.1% 的平均迁移结果指数,在 47 个仓库上较直接翻译取得了 17 至 59 个百分点的结果指数优势。这些结果表明,可检查的语义瓶颈可提升仓库规模代码移植的覆盖范围与集成度,同时可为多数仓库实现更低成本的代码生成。
英文摘要
Legacy software repositories embed decades of domain knowledge in undocumented code, making understanding and modernization difficult. We treat a program as the implementation of an unobserved, declarative description of its computation and investigate whether making this latent declarative representation explicit improves repository-scale porting. ADFD-Migrate approximates the latent representation with an annotated data-flow diagram (ADFD) of processes, data stores, external entities, flows, and behavioral contracts. An LLM infers the source ADFD from bounded repository context, guided by static-analysis coverage checks. Dependency-aware chunking orders bounded process groups for target-language generation. Differences between the source ADFD and a statically recovered target ADFD then guide regeneration. We evaluate ADFD-Migrate on f2x50, a new benchmark of 50~Fortran repositories spanning 1.5k--1.6M lines of code and three complexity tiers, and assess the resulting ports along two dimensions: porting soundness, measured by source-oracle behavioral agreement, and porting completeness, measured by a composite migration outcome index. Against 382 curated Fortran-oracle probes, the generated Python passes 327 (85.6\%), with 40 repositories passing every attempted probe. ADFD-Migrate exposes all 382 planned behaviors as runnable targets, compared with 99 and 98 for direct and repository-context translation and 69 and 30 for the static-profile and dependency-chunking ablations. It also achieves a 93.1\% mean migration outcome index and a 17--59 percentage-point outcome-index advantage over direct translation on 47 repositories. These results suggest that an inspectable semantic bottleneck can improve the coverage and integration of repository-scale migration while enabling lower-cost generation for many repositories.
发表机构
- BITS Pilani(BITS Pilani( Pilani 理工学院))
- UNSW Sydney(新南威尔士大学悉尼分校)
机构由 AI 辅助整理,请以论文原文为准。