arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DepWareTrans:跨可共执行语言的依赖感知增量仓库迁移

DepWareTrans: Dependency-Aware Incremental Repository Migration across Co-executable Languages

Sivajeet Chand, Alexander Pretschner, Steve Haupt, Derui Zhu, Sushant Kumar Pandey

arXiv 2608.14128首次发表:更新:

AI 中文总结

提出依赖感知增量迁移框架,构建依赖图分组文件进行批量翻译,在 STAR 仓库等上实现 100% 编译和测试成功率,提升仓库级代码翻译的可扩展性与可靠性。

AI 中文摘要

仓库级代码翻译对于现代化遗留系统至关重要,但现有基于大语言模型(LLM)的方法以文件为单位操作,无法扩展到具有复杂文件间依赖关系的代码库中。这一局限在我们的工业场景中十分明显:我们计划将生产仓库 STAR 从 Java 迁移到 Kotlin,但基于文件的方法会产生碎片化结果,无法实现端到端正确性。在本文中,我们表明仓库级迁移失败的主要原因是依赖不一致。通过对开源系统和工业系统的实证研究,我们发现大多数错误源于未解决的跨文件依赖,仅靠迭代反馈无法有效处理这些依赖。我们提出了一种依赖感知增量迁移框架,将翻译单位从单个文件提升为依赖一致的批次。我们的方法构建依赖图,将相互依赖的文件分组,并进行基于编译和测试驱动验证的批量翻译。我们在一个 5.1 万行代码(LOC)的工业系统及多个跨可互操作语言对(Java-Kotlin、Java-Scala 和 C#-F#)的仓库上评估了我们的方法。在 STAR 仓库上,基于文件的方法实现了 38.16% 的编译成功率和 9.39% 的测试成功率,而我们的方法在评估设置下实现了 100% 的编译和测试成功率,且在少量迭代内收敛。这些结果表明,依赖感知批次处理提高了仓库级代码翻译的可扩展性和可靠性。

英文摘要

Repository-level code translation is critical for modernizing legacy systems, yet existing approaches based on large language models (LLMs) operate at the file level and fail to scale to codebases with complex inter-file dependencies. This limitation is evident in our industrial setting, where we aim to migrate a production repository (STAR) from Java to Kotlin, but file-level approaches produce fragmented results and fail to achieve end-to-end correctness. In this paper, we show that the primary cause of failure at the repository level is dependency inconsistency. Through an empirical study on open-source and industrial systems, we find that most errors arise from unresolved cross-file dependencies that cannot be effectively addressed by iterative feedback alone. We propose a dependency-aware incremental migration framework that elevates the unit of translation from individual files to dependency-consistent batches. Our approach constructs a dependency graph, groups interdependent files, and performs batched translation with iterative compile- and test-driven validation. We evaluate our method on a 51K line of code (LOC) industrial system and multiple repositories across interoperable language pairs (Java-Kotlin, Java-Scala, and C#-F#). On the STAR repository, file-level approaches achieve 38.16% compilation and 9.39% test success, whereas our approach achieves 100% compilation and test success across the evaluated settings, converging within a small number of iterations. These results show that dependency-aware batching improves scalability and reliability in repository-level code translation.

CommentsAccepted for publication in the Industry Showcase Track of the 41st IEEE/ACM International Conference on Automated Software Engineering, which will take place in Munich, Germany during October 12-16, 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑