arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.20257cs.SE

CodeTransBenchmark:评估基于LLM的跨编程语言代码翻译与修复

CodeTransBenchmark: Evaluating LLM-based Code Translation and Repair Across Programming Languages

Vera Kowalczuk, Oliver Weißl, Severin Kacianka, Andrea Stocco

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出CodeTransBenchmark框架,评估八个LLM在三个数据集和12个语言对上的代码翻译与修复能力,发现通用模型受目标语言语法限制,而迭代修复可显著提升准确性。

中文摘要 AI 辅助

在广泛文本和代码语料库上预训练的大型语言模型(LLMs)展现了有前景的代码生成能力,并在代码翻译领域引起了越来越多的关注。在本工作中,我们研究了LLMs在代码翻译和翻译错误修复中的有效性。首先,我们提出了CodeTransBenchmark,一个用于评估基于LLM的翻译和修复的框架,并设计了一种后处理策略,以从不一致的LLM输出中提取代码。然后,我们讨论了一项实证研究,该研究在三个数据集和12个语言对上评估了八个模型,其中我们按错误对不正确的翻译进行分类,以识别现有LLMs的弱点。我们的工作表明,虽然专门针对多语言编码训练的LLMs(如Codestral)能正确翻译大部分代码,但大多数通用模型在目标语言的语法规则上存在困难。对错误翻译的分析揭示了所涉及编程语言与训练数据之间的相互关系对有效性的重大影响。我们表明,一种通用的后处理方法必须容忍不一致性,并利用LLM答案的可预测性。此外,我们表明,通过自动化反馈进行迭代翻译修复能显著提高翻译准确性。虽然我们的综合发现突出了LLMs自动化代码翻译的潜力,但在实践中有效部署基于LLM的代码翻译将需要具有更大上下文窗口的模型。

英文摘要

Large Language Models (LLMs) pre-trained on expansive text and code corpora have revealed promising code generation abilities and have attracted increasing attention in code translation. In this work, we investigate the effectiveness of LLMs in code translation and translation error repair. First, we present CodeTransBenchmark, a framework for evaluating LLM-based translation and repair and devise a post-processing strategy to extract code from inconsistent LLM outputs. Then, we discuss an empirical study evaluates eight models on three datasets and 12 language pairs, in which we categorize incorrect translations by errors to identify weaknesses of existing LLMs. Our work shows that while LLMs specifically trained for multi-lingual coding, like Codestral, correctly translate the majority of code, most general-purpose models struggle with the syntactic rules of the target language. The analysis of erroneous translations reveals the substantial impact of the interrelationship between involved programming languages and training data on the effectiveness. We show that a general post-processing approach must tolerate inconsistencies and leverage the predictability of LLM answers. Further, we show that iterative translation repair via automated feedback significantly improves translation accuracy. While our combined findings highlight the potential of LLMs to automate code translation, an effective deployment of LLM-based code translation in practice would require models with larger context windows.

发表机构

  • Technical University of Munich(慕尼黑工业大学)
  • fortiss GmbH(fortiss有限公司)

机构由 AI 辅助整理,请以论文原文为准。

↑