arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.19966cs.SE

通过规则引导推理和强化学习实现可靠的C到Rust翻译

Towards Reliable C-to-Rust Translation with Rule-Guided Reasoning and Reinforcement Learning

Feng Luo, Jiachen Liu, Cuiyun Gao, Jia Feng, Kui Liu

首次发表
浏览论文内容

中文总结 AI 辅助

研究旨在解决利用大语言模型进行C到Rust自动翻译时存在的问题,提出TRAVEL框架,通过规则引导推理和强化学习两个模块,在多个数据集上评估,该框架提升了翻译准确率、成功率并降低不安全率。

中文摘要 AI 辅助

将遗留C程序迁移到Rust已成为提高软件内存安全性同时降低手动重写高成本的重要方向。利用大语言模型(LLMs)进行自动C到Rust翻译是一个有前景的方向。然而,现有基于LLMs的方法存在局限。一方面,LLMs识别Rust特定规则的能力有限,Rust语法处理不当常导致翻译错误;另一方面,现有LLMs难以准确捕捉复杂代码语义,也会导致翻译错误。为应对这些挑战,我们提出了一个通过规则引导推理和强化学习的翻译框架TRAVEL,它由两个模块组成。第一个模块采用基于蒙特卡洛树搜索(MCTS)的推理路径构建,由Rust特定规则引导,引导搜索朝着尊重LLMs常违反的语法规则的翻译步骤进行。第二个模块引入强化学习,将执行反馈与推理质量信号相结合,鼓励模型构建准确捕捉程序语义的推理路径,从而确保生成的Rust代码保留原始C程序的预期行为。我们在三个数据集上评估了TRAVEL:xCodeEval(一个公共基准)、OS - Bench(从Linux内核收集的函数)和HW - Bench(华为的一个工业数据集)。在xCodeEval上,TRAVEL在三个主干LLMs上均优于所有基线。特别是,与最强的提示基线IRENE相比,TRAVEL将计算准确率(CA)提高了26.22%,编译成功率(CSR)提高了18.77%。在HW - Bench和OS - Bench上,TRAVEL分别将CSR进一步提高了18.28%和16.51%,同时分别将不安全率(UR)降低了13.06%和13.08%。

英文摘要

The migration of legacy C programs to Rust has become an important direction for improving software memory safety while alleviating the high cost of manual rewriting. Leveraging large language models (LLMs) for automated C-to-Rust translation has emerged as a promising direction. However, existing LLM-based approaches remain limited. On the one hand, LLMs exhibit limited capability in identifying Rust-specific rules, and inadequate handling of Rust syntax often results in incorrect translations. On the other hand, existing LLMs often struggle to accurately capture the semantics of complex code, resulting in incorrect translations. To address these challenges, we propose a Translation fRAmework Via rule-guided reasoning and rEinforcement Learning, namely TRAVEL, consisting of two modules. The first module employs Monte Carlo Tree Search (MCTS)-based reasoning path construction guided by Rust-specific rules, steering the search toward translation steps that respect the syntactic rules that LLMs frequently violate. The second module introduces reinforcement learning that couples execution feedback with reasoning-quality signals, encouraging the model to construct reasoning paths that accurately capture program semantics, thereby ensuring that the generated Rust code preserves the intended behavior of the original C program. We evaluate TRAVEL on three datasets: xCodeEval (a public benchmark), OS-Bench (functions collected from the Linux kernel), and HW-Bench (an industrial dataset from Huawei). On xCodeEval, TRAVEL outperforms all baselines across three backbone LLMs. In particular, compared to the strongest prompting baseline IRENE, TRAVEL improves computational accuracy (CA) by 26.22% and compilation success rate (CSR) by 18.77%. On HW-Bench and OS-Bench, TRAVEL further improves CSR by 18.28% and 16.51%, respectively, while reducing unsafe rate (UR) by 13.06% and 13.08%, respectively.

补充信息

↑