微调Qwen3-27B实现C到Rust代码翻译:预训练、调试感知监督微调与任务特定监督微调的三阶段课程
Fine-Tuning Qwen3-27B for C-to-Rust Code Translation: A Three-Stage Curriculum of Pretraining, Debugging-Aware SFT, and Task-Specific SFT
浏览论文内容
中文总结 AI 辅助
本研究针对C到Rust代码翻译任务,对Qwen3-27B采用三阶段微调课程,结合SACTOR框架评估后,其性能优于基线模型及其他LLM。
中文摘要 AI 辅助
将C代码转换为安全、符合习惯用法的Rust代码是一项长期存在的软件工程目标,因为这可以在保留遗留系统功能行为的同时,消除整类内存安全漏洞。大语言模型(LLM)在该任务上展现出潜力,但直接应用时通常表现不佳,因为通用预训练很少强调符合习惯用法的Rust生成、跨语言语义等价性,或推理及修复编译器/运行时反馈的能力。本报告描述了应用于Qwen3-27B的三阶段微调课程,旨在逐步将模型专门化以完成C到Rust(C2Rust)翻译任务:(1)在以Rust为中心的语料库上进行持续预训练,以增强模型对符合习惯用法的Rust语法及标准库使用的先验知识;(2)在microsoft/Verus_Training_Data数据集上进行监督微调(SFT),以灌输Rust代码的调试与自我修复行为;(3)在源自LeetCode问题的配对C/Rust解决方案上进行任务特定SFT,以教授直接语义翻译。我们使用SACTOR的智能体、静态分析引导验证框架评估所得模型,该框架执行结构感知的两阶段(非习惯用法到习惯用法)翻译,并进行基于外部函数接口(FFI)的端到端(E2E)测试。我们报告了成功率、习惯用法性(Clippy lint数量、不安全代码占比)及失败模式分析,并在相同框架下将我们的微调模型与基线Qwen3-27B及其他LLM进行比较。
英文摘要
Translating C code into safe, idiomatic Rust is a longstanding software-engineering goal because it can eliminate entire classes of memory-safety vulnerabilities while preserving the functional behavior of legacy systems. Large language models (LLMs) have shown promise for this task but typically underperform when applied off-the-shelf, since general-purpose pretraining rarely emphasizes idiomatic Rust generation, cross-language semantic equivalence, or the ability to reason about and repair compiler/runtime feedback. In this report we describe a three-stage fine-tuning curriculum applied to Qwen3-27B that is designed to progressively specialize the model for the C-to-Rust (C2Rust) translation task: (1) continued pretraining on Rust-centric corpora to strengthen the model's prior over idiomatic Rust syntax and standard-library usage; (2) supervised fine-tuning (SFT) on the microsoft/Verus_Training_Data dataset to instill debugging and self-repair behavior over Rust code; and (3) task-specific SFT on paired C/Rust solutions derived from LeetCode problems to teach direct semantic translation. We evaluate the resulting model using the agentic, static-analysis-guided verification framework of SACTOR, which performs structure-aware, two-phase (unidiomatic to idiomatic) translation with foreign-function-interface (FFI)-based end-to-end (E2E) testing. We report success rate, idiomaticity (Clippy lint counts, unsafe-code fraction), and failure-mode analyses, and compare our fine-tuned model against baseline Qwen3-27B and other LLMs evaluated under the same framework.
发表机构
- Northeastern University(东北大学)
- EmbodyX Inc(EmbodyX公司)
- Aibao LLC(Aibao有限责任公司)
机构由 AI 辅助整理,请以论文原文为准。