基于执行可验证监督的强化学习引导小众多语言代码翻译
Bootstrapping Niche Multilingual Code Translation via Reinforcement Learning with Execution-Based Verifiable Supervision
浏览论文内容
中文总结 AI 辅助
本研究针对多语言代码翻译中热门语言外场景并行监督稀缺、翻译不可执行的问题,提出基于执行监督的偏好强化学习方法,引入HumanEval-X++基准,用Qwen模型验证了方法的有效性。
中文摘要 AI 辅助
代码翻译必须在多种编程语言间保留可执行行为,但神经代码翻译大多聚焦于C++、Java和Python等少数热门语言,这导致在并行监督稀缺的小众多对多场景中,会生成看似合理但不可执行的翻译结果。我们提出一种基于执行监督的偏好强化学习方法解决该场景问题:首先将可验证的种子Python程序扩展为经执行验证的多语言代码池;利用该代码池,基础大语言模型(LLM)生成跨语言对的翻译候选,再根据执行结果为候选打上标签;将得到的偏好用于训练奖励模型,以评估跨语言翻译质量;最后使用GRPO算法,以奖励模型为信号,在600个定向语言对(25种源语言×24种目标语言)上优化基础LLM。为评估小众翻译能力,我们引入HumanEval-X++,这是一种基于执行的基准,将HumanEval-X扩展至广泛的多对多语言空间。我们使用Qwen-3.5 4B和9B模型评估所提方法,在HumanEval-X++及现有基准上,其相比未训练基线取得了持续提升;其中4B模型在HumanEval-X++的所有语言上平均提升13%,在中端语言上提升21%。本研究确立了一套可靠的数据生成、训练及基准测试方法,为进一步提升编程语言多对多翻译的质量奠定了基础。
英文摘要
Code translation must preserve executable behavior across many programming languages, yet neural code translation has largely focused on a few popular languages such as C++, Java, and Python. This leaves a niche, many-to-many setting where parallel supervision is sparse, producing plausible but non-executable translations. We address this setting with preference-based reinforcement learning driven by execution-based supervision. Our pipeline firstly expands verifiable seed Python programs into a multilingual pool of execution-validated codes. Using the pool, a base LLM generates translation candidates across language pairs, which we label by their execution outcomes. The resulting preferences are used to train a reward model that scores cross-language translation quality. Finally, we optimize our base LLMs with GRPO over 600 directed language pairs (25 x 24) using the reward model as a signal. To evaluate the niche translation capability, we introduce HumanEval-X++, an execution-based benchmark that extends HumanEval-X to a broad many-to-many language space. We evaluate our approach using Qwen-3.5 4B and 9B models. On HumanEval-X++ and existing benchmarks, it yields consistent gains over the untrained baselines. In particular, the 4B model achieves an average improvement of 13% across all languages on HumanEval-X++, with a gain of 21% on mid-tier languages. Our study establishes a reliable approach of data generation, training, and benchmarking, paving the way toward further bootstrapping the quality of many-to-many translation for programming languages.