AI 中文总结
针对现实世界噪声指令对代码LLM的鲁棒性挑战,提出NTRL-Code框架,在测试阶段利用无标签噪声数据,通过保守自去噪和AST结构聚合实现鲁棒自我进化,在三个基准上取得稳健改进。
AI 中文摘要
大语言模型(LLMs)在各种代码相关任务中展现出了卓越的性能。然而,与通常高质量且无错误的人工精心策划的数据集不同,现实世界中的用户指令往往模糊且容易出错,这给代码大语言模型的鲁棒性带来了重大挑战。此外,面向鲁棒性的微调依赖于配对的干净-噪声样本,这些样本的策划成本高昂,并且需要复杂的噪声模拟技术。为了应对这些挑战,我们提出了噪声测试时强化学习框架(NTRL-Code),该框架仅使用测试阶段的无标签噪声数据即可实现代码大语言模型的鲁棒自我进化。具体来说,NTRL-Code使用保守的自去噪来获得更干净的语义锚点用于目标估计,并采用基于抽象语法树(AST)的结构聚合机制从多个候选程序中估计代理目标。随后,策略在原始噪声提示上使用结合格式有效性、代码相似性和抗重复信号的混合奖励进行优化。在三个基准上的大量实验(每个基准包含字符级、词级和段落级扰动)表明,NTRL-Code产生了稳健且一致的改进,稳定了各种基础模型的预测。我们的代码可在以下网址获取:此https URL。
英文摘要
Large language models (LLMs) have demonstrated remarkable performance across various code-related tasks. However, unlike carefully curated datasets that are typically high-quality and error-free, real-world user instructions are often vague and error-prone, posing significant challenges to the robustness of code LLMs. Furthermore, robustness-oriented fine-tuning relies on paired clean-noisy samples, which are costly to curate and require sophisticated noisy simulation techniques. To address these challenges, we propose the Noisy Test-time Reinforcement Learning framework (NTRL-Code), which enables robust self-evolution of code LLMs using only unlabeled noisy data during the testing stage. Specifically, NTRL-Code uses conservative self-denoising to obtain a cleaner semantic anchor for target estimation, and employs an abstract-syntax-tree (AST)-based structural aggregation mechanism to estimate a proxy target from multiple candidate programs. The policy is then optimized on the original noisy prompts with a hybrid reward that combines format validity, code similarity, and anti-repetition signals. Extensive experiments on three benchmarks, each incorporating character-level, word-level, and paragraph-level perturbations, demonstrate that NTRL-Code yields robust and consistent improvements, stabilizing the predictions of various base models. Our code is available at https://github.com/Xikai97/NTRL-Code.
CommentsThis paper has been accepted by EMNLP 2026