代码翻译技术的大规模实证研究
An Extensive Empirical Study on Code Translation Technique
AI总结:
本研究通过大规模实证对比,发现LLM及基于LLM的代码翻译方法在方法级正确性上优于学习类方法,类级翻译更难,静态与逻辑错误是主要挑战,为代码翻译技术开发提供指导。
AI中文摘要:
自动化代码翻译对软件演化愈发重要,但基于学习的技术与基于大语言模型(LLM)技术的相对优势和局限性仍未得到充分理解。为填补这一空白,我们开展大规模实证研究,对比不同方法学范式与翻译粒度下的代表性代码翻译技术。我们在涉及多种编程语言的多语言方法级和类级基准上,评估基于学习的方法、基于LLM的方法以及通用LLM。我们的分析考虑可执行正确性、代码相似度、翻译方向、翻译粒度和失败模式。结果表明,在方法级正确性方面,LLM和基于LLM的方法总体优于基于学习的方法,尽管仅靠相似度指标无法可靠反映功能正确性。翻译方向对性能影响显著,尤其在类型系统特性不同的语言间翻译时。类级翻译比方法级翻译困难得多,因为它需要保留全局语义、接口、成员关系和跨方法依赖。错误分析进一步显示,静态语义错误和逻辑错误是现有代码翻译系统的主要挑战。这些发现为开发更鲁棒、感知类型、感知结构和感知上下文的代码翻译技术提供了实证依据和实践指导。
英文摘要:
Automated code translation is increasingly important for software evolution, yet the relative strengths and limitations of learning-based and large language model (LLM)-based techniques remain insufficiently understood. To address this gap, we conduct a large-scale empirical study comparing representative code translation techniques across methodological paradigms and translation granularities. We evaluate learning-based methods, LLM-based methods, and general-purpose LLMs on multilingual method-level and class-level benchmarks involving multiple programming languages. Our analysis considers executable correctness, code similarity, translation direction, translation granularity, and failure patterns. The results show that LLMs and LLM-based methods generally outperform learning-based methods in method-level correctness, although similarity metrics alone do not reliably reflect functional correctness. Translation direction substantially affects performance, particularly when translating between languages with different type-system characteristics. Class-level translation remains considerably more difficult than method-level translation because it requires preserving global semantics, interfaces, member relationships, and cross-method dependencies. Our error analysis further shows that static semantic errors and logical errors are the primary challenges in existing code translation systems. These findings provide empirical evidence and practical guidance for developing more robust, type-aware, structure-aware, and context-aware code translation techniques.