用于在编程竞赛中取得金牌级表现的训练后语言模型
Post-Training Language Models for Gold-Medal Performance in Coding Competitions
浏览论文内容
中文总结 AI 辅助
本研究提出结合SFT、RL及GenCorrect策略的训练后语言模型,在编程竞赛中超越人类选手,成为首个在IOI题集得分超最高人类参赛者的AI系统。
中文摘要 AI 辅助
竞赛编程已成为测试大型语言模型推理能力的关键指标,国际信息学奥林匹克竞赛(IOI)和国际大学生程序设计竞赛(ICPC)等赛事代表了该领域最具挑战性的场景。我们提出了一种端到端的专业化流程,结合大规模问题筛选、合成推理轨迹、监督微调(SFT)和强化学习(RL)。利用22000个筛选后的问题,我们通过SFT和RL训练了Nemotron-3-Nano-CC(30B-A3B),并仅通过SFT训练了Nemotron-3-Ultra-CC(550B-A55B)。我们还引入了GenCorrect,一种反馈驱动的测试时计算策略,可迭代生成、评估和优化多样化的解决方案。在2025年IOI赛事中,Nano-CC在训练后得分从130提升至291,结合GenCorrect后得分达468,超过了438.3的金牌阈值,而Ultra-CC得分达到502。基于这些结果,我们开发了一个竞赛专用的Ultra-CC系统,并在2026年IOI赛事中进行了前瞻性评估。在与人类参赛者相同的时间、联网权限和提交约束下,该系统得分535.4(满分600),超过了361.12的金牌阈值和498.27的人类最高分。据我们所知,这是第一个在IOI题集上得分超过最高人类参赛者的AI系统。
英文摘要
Competitive programming has become a key test of large language model reasoning, with international competitions such as IOI and ICPC representing its most challenging settings. We present an end-to-end specialization pipeline combining large-scale problem curation, synthetic reasoning traces, supervised fine-tuning (SFT), and reinforcement learning (RL). Using 22,000 curated problems, we train Nemotron-3-Nano-CC (30B-A3B) with SFT and RL and Nemotron-3-Ultra-CC (550B-A55B) with SFT alone. We further introduce GenCorrect, a feedback-driven test-time compute strategy that iteratively generates, evaluates, and refines diverse solutions. On IOI 2025, Nano-CC improves from 130 points to 291 after post-training and to 468 with GenCorrect, exceeding the gold threshold of 438.3 while Ultra-CC reaches 502. Guided by these results, we develop a competition-specific Ultra-CC system and evaluate it prospectively during IOI 2026. Under the same time, internet-access, and submission constraints as human contestants, it scores 535.4 out of 600, exceeding both the gold threshold of 361.12 and the top human score of 498.27. To our knowledge, this is the first AI system to outscore the highest-scoring human contestant on an IOI problem set.
发表机构
- NVIDIA(英伟达)
机构由 AI 辅助整理,请以论文原文为准。