通过逐定理符号验证器生成反例:当模仿有害而强化修复
Counterexample Generation via Per-Theorem Symbolic Verifiers: When Imitation Hurts and Reinforcement Repairs
- National School of Artificial Intelligence(国家人工智能学院)
- University of Birmingham(伯明翰大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对大模型证伪差距,提出基于逐定理验证器的反例生成框架SymCE,发现模仿SFT有害而强化学习RLVR修复并提升性能,4B模型超越7B专家并媲美商业API。
AI中文摘要:
大型语言模型常常能正向证明一个定理,却无法推翻一个密切相关的错误命题:这种证伪差距是监督微调无法弥合甚至可能主动加剧的。我们将反例生成框架化为针对确定性逐定理Python验证器的约束见证生成,并发布了SymCE,一个包含4,707个错误本科代数与实分析猜想的数据集,每个猜想都配有可执行的验证器。该验证器同时充当奖励函数,使SymCE成为一个训练环境。在此预言机下,使用SFT后接GRPO训练Qwen3-4B,揭示了一个模仿陷阱:仅使用反例的SFT将真定理识别率从0.27降至0.00,而使用稀疏仅结果奖励的RLVR修复了这一问题并超越基线,达到0.66。这种崩溃在四个种子及Gemma-3-4B上重复出现。稀疏与密集奖励在域内成功率上统计上无显著差异,但在留出校准探针上相差33个百分点,我们将这种分离归因于部分学分项。我们的4B模型优于所有评估的7B开源数学专家模型,与六个前沿商业API保持竞争力,并在不改变提示的情况下迁移到GSM8K、MATH-500和MMLU大学数学。对177个验证器决策的人工审计发现97.7%的准确率。代码、数据、验证器模块和注释:此https URL。
英文摘要:
Large language models often solve a theorem forward yet fail to disprove a closely related false one: a falsification gap that supervised fine-tuning does not close and can actively worsen. We frame counterexample generation as constrained witness emission against a deterministic per-theorem Python verifier, and release SymCE, a corpus of 4,707 false undergraduate-algebra and real-analysis conjectures, each paired with executable verifiers. The verifier also serves as the reward function, making SymCE a training environment. Training Qwen3-4B with SFT followed by GRPO under this oracle reveals an imitation trap: counterexample-only SFT collapses true-theorem recognition from 0.27 to 0.00, while RLVR with a sparse outcome-only reward repairs this and exceeds the base, to 0.66. The collapse replicates across four seeds and on Gemma-3-4B. Sparse and dense rewards yield statistically indistinguishable in-domain success yet diverge by 33 points on a held-out calibration probe, a dissociation we trace to the partial-credit term. Our 4B model outperforms every evaluated 7B open-weights math specialist, remains competitive with six frontier commercial APIs, and transfers under unchanged prompting to GSM8K, MATH-500 and MMLU-college-math. A human audit of 177 verifier decisions finds 97.7% accuracy. Code, data, verifier modules and annotations: https://github.com/ce-rlvr/SymCE.