AI 中文总结
本文提出一种开放配方,通过后训练和测试时计算流水线,利用Nemotron 3 Ultra生成自然语言证明,在IMO 2026获30/42分达金牌线,并开源模型、数据、代码和基准。
AI 中文摘要
我们研究了模型后训练和测试时推理设计如何影响硬奥林匹克数学的自然语言证明生成。从Nemotron 3 Ultra出发,我们使用监督微调和强化学习训练了两个专家检查点,并评估了检查点选择、验证和细化。基于这些发现,我们提出了一个开放模型的测试时计算流水线。该系统完全以自然语言运行,没有形式化证明器、外部工具或互联网访问。三个Nemotron 3 Ultra检查点——通用可用模型和两个后训练专家——驱动一个迭代搜索,生成、验证和细化候选证明;一个独立的高计算阶段随后选择每个最终提交。该系统在IMO 2026中获得了42分中的30分,达到了金牌门槛。我们发布了两个后训练检查点以及训练数据、训练和推理代码、提交的解决方案,以及Nemotron-IMO-Bench,一个包含200个新颖奥林匹克级问题的新基准。
英文摘要
We study how model post-training and test-time inference design affect natural-language proof generation for hard olympiad mathematics. Starting from Nemotron 3 Ultra, we train two specialist checkpoints using supervised fine-tuning and reinforcement learning, and evaluate checkpoint choice, verification, and refinement. Based on these findings, we present an open-model test-time-compute pipeline. The system operates entirely in natural language, with no formal prover, external tools, or internet access. Three Nemotron 3 Ultra checkpoints - the general-availability model and two post-trained specialists - power an iterative search that generates, verifies, and refines candidate proofs; a separate high-compute stage then selects each final submission. The system scored 30 out of 42 points at IMO 2026, reaching the gold-medal threshold. We release the two post-trained checkpoints as well as the training data, the training and inference code, the submitted solutions, and Nemotron-IMO-Bench, a new benchmark of 200 novel olympiad-level problems.