AI 中文总结
该研究提出形式化翻译器MerC,结合MerC与大语言模型(LLMs)的搭档方法,在宏翻译任务中比单独使用任一技术更优,可降低失败率并提升翻译覆盖率。
AI 中文摘要
现代关键软件基础设施大多用C语言编写。由于C语言缺乏内存安全性,研究人员正在探索将C语言自动翻译为Rust等更安全的语言。但现实世界中的C软件不止包含C代码,还常使用名为宏(macros)的命名代码片段,这些片段并非C语言本身的组成部分。现有最先进的技术会先对C代码进行预处理再翻译,以此避免翻译宏,但该方法生成的译文与原始C代码差异较大,因为预处理会内联所有宏定义。为在译文中保留宏的使用,我们研究了宏与C语言共有的语言特征,并将其提炼为首个形式化规范的翻译器MerC。为评估MerC,我们推出了首个宏翻译基准MacroBench,其测试用例基于从现实世界C程序中随机采样的宏。研究发现,MerC支持MacroBench中50%的宏测试用例。我们还使用MacroBench评估大语言模型(LLMs)在之前未被研究过的宏翻译任务中的表现:LLMs对MacroBench的翻译量比MerC多22%至77%,但其中8%和28%的译文存在错误,需要开发者进一步验证;相比之下,MerC生成的译文全部正确。我们的关键见解是:先运行MerC,再对剩余部分使用LLMs,这种搭档方法比单独使用任一技术能获得更大收益——其平均失败率比LLMs低32%,同时比MerC多翻译平均51%的测试用例。
英文摘要
Modern critical software infrastructure is largely written in C. Since C lacks memory safety, researchers are investigating automatic translation of C to safer languages like Rust. But real-world C software consists of more than just C code, often using named code fragments called macros which are not part of the C language proper. State-of-the-art techniques avoid translating macros by preprocessing C code first before translating it. But this approach produces translations that are dissimilar to the original C code, because preprocessing inlines all macro definitions. To preserve macro usage in translated code, we study the language features that macros and C share and distill them into the first formally-specified translator, MerC. To evaluate MerC, we introduce the first macro translation benchmark, MacroBench, with test cases based on macros randomly sampled from real-world C programs. We find that MerC supports 50% of MacroBench's macro test cases. We also use MacroBench to evaluate how effective large language models (LLMs) are at performing the previously-unstudied task of macro translation. LLMs translate 22% to 77% more of MacroBench than MerC, but with 8% and 28% of these translations being incorrect translations requiring additional validation by developers. In contrast, MerC only produces correct translations. Our key insight is that running MerC first then using LLMs on the remainder reaps greater benefits than using either technique alone. This tag team approach has an average failure rate 32% lower than that of LLMs, while also translating an average of 51% more test cases than MerC.
Comments13 pages, 7 figures, 2 tables, 8 listings, conference. To be published in ASE 2026