RepoTransBench: 一个面向真实世界的多语言代码仓库翻译基准
RepoTransBench: A Real-World Multilingual Benchmark for Repository-Level Code Translation
- Sun Yat-sen University, China(中山大学)
- Monash University, Australia(墨尔本大学)
- Huawei Cloud Computing Technologies Co., Ltd, China(华为云计算技术有限公司)
- Chongqing University, China(重庆大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
RepoTransBench提出一个现实世界多语言代码仓库翻译基准,评估代码仓库级翻译的挑战,揭示不同语言对方向的翻译难度差异。
AI中文摘要:
代码仓库级别的翻译指的是将整个代码仓库从一种编程语言翻译成另一种语言,同时保持源仓库的功能。已经提出了许多基准来评估此类代码翻译器的性能。然而,之前的基准大多提供细粒度的样本,只关注代码片段、函数或文件级别的代码翻译。这些基准不能准确反映真实世界的需求,其中整个仓库往往需要翻译,涉及更长的代码长度和更复杂的功能。为了解决这一缺口,我们提出了一个新的基准,名为RepoTransBench,这是一个具有13种语言对的现实世界多语言代码仓库翻译基准,包含1897个现实世界仓库样本,并具有自动可执行的测试套件。此外,我们引入了RepoTransAgent,一个通用代理框架,用于执行代码仓库级别的翻译。我们使用几种方法和基础LLMs评估了该基准的挑战和代理的有效性,发现代码仓库级别的翻译仍然具有挑战性,其中表现最好的方法只实现了32.8%的成功率。此外,我们的分析揭示了翻译难度在语言对方向上存在显著差异,动态到静态语言的翻译比反向方向(静态到动态)更具挑战性(动态到静态翻译低于10%,而静态到动态翻译为45-63%)。最后,我们进行了详细的错误分析,并指出现有LLMs在代码仓库级别翻译中的不足,这可能为进一步改进提供参考。我们提供代码和数据在https://github.com/DeepSoftwareAnalytics/RepoTransBench。
英文摘要:
Repository-level code translation refers to translating an entire code repository from one programming language to another while preserving the functionality of the source repository. Many benchmarks have been proposed to evaluate the performance of such code translators. However, previous benchmarks mostly provide fine-grained samples, focusing at either code snippet, function, or file-level code translation. Such benchmarks do not accurately reflect real-world demands, where entire repositories often need to be translated, involving longer code length and more complex functionalities. To address this gap, we propose a new benchmark, named RepoTransBench, which is a real-world multilingual repository-level code translation benchmark featuring 1,897 real-world repository samples across 13 language pairs with automatically executable test suites. Besides, we introduce RepoTransAgent, a general agent framework to perform repository-level code translation. We evaluate both our benchmark's challenges and agent's effectiveness using several methods and backbone LLMs, revealing that repository-level translation remains challenging, where the best-performing method achieves only a 32.8% success rate. Furthermore, our analysis reveals that translation difficulty varies significantly by language pair direction, with dynamic-to-static language translation being much more challenging than the reverse direction (achieving below 10% vs. static-to-dynamic at 45-63%). Finally, we conduct a detailed error analysis and highlight current LLMs' deficiencies in repository-level code translation, which could provide a reference for further improvements. We provide the code and data at https://github.com/DeepSoftwareAnalytics/RepoTransBench.