从推理字符串到偏序:通过商策略优化的验证器认证规则迁移
From Reasoning Strings to Partial Orders: Verifier-Certified Rule Transport through Quotient Policy Optimization
AI总结:
本文提出验证器认证规则迁移(VCRT),通过反菱形监督保留真实先决条件,在多个推理环境中将宏通过率提升13.06个百分点,但未证实精确轨道聚合的普遍优势。
AI中文摘要:
许多计算由于独立子目标或不相交的状态更新可以交换,因此允许多种有效的执行顺序。使用可验证奖励的强化学习通常将每个成功轨迹视为独立的令牌序列,因此序列化选择可能被误认为是逻辑依赖。我们引入了验证器认证规则迁移(VCRT),它使用原生验证器重放相邻操作对。两个顺序都被接受并达到相同规范状态的对提供交换证书;被拒绝或改变状态的反转提供反菱形。VCRT 使用反菱形来保留真正的先决条件,并将策略信用分配给每个认证轨道的总概率质量。它还约束了交换后一致性、源保留和策略漂移。我们通过共享的匿名关系图接口,在 ProofWriter、CLRS 和 Lean 上评估了留一环境迁移。所有训练和检查点决策在保留评估之前冻结,保留评估使用每个项目的单条贪婪轨迹,无需搜索或验证器反馈。VCRT 获得了 77.60% 的宏通过率,而最强匹配基线为 64.53%,配对增益为 13.06 个百分点(95% 自助置信区间 [12.58, 13.54])。Lean 贡献了大部分增益,为 33.49 个百分点,而 ProofWriter 和 CLRS 平均提高了 2.85 个百分点。机制测试一致支持反菱形监督,而 No-Orbit 与完整 VCRT 在统计上无显著差异。证据并未确立精确轨道聚合的普遍益处。
英文摘要:
Many computations admit several valid execution orders because independent subgoals or disjoint state updates can commute. Reinforcement learning with verifiable rewards usually treats each successful trace as a separate token sequence, so serialization choices can be mistaken for logical dependencies. We introduce Verifier-Certified Rule Transport (VCRT), which replays adjacent operation pairs with native verifiers. Pairs whose two orders are accepted and reach the same canonical state provide commutation certificates; rejected or state-changing reversals provide anti-diamonds. VCRT uses anti-diamonds to preserve genuine prerequisites and assigns policy credit to the total probability mass of each certified orbit. It also constrains post-swap consistency, source retention, and policy drift. We evaluate leave-one-environment-out transfer across ProofWriter, CLRS, and Lean through a shared anonymized relation-graph interface. All training and checkpoint decisions are frozen before held-out evaluation, which uses one greedy trajectory per item without search or verifier feedback. VCRT obtains a 77.60% macro pass rate versus 64.53% for the strongest matched baseline, a paired gain of 13.06 points (95% bootstrap CI [12.58, 13.54]). Lean accounts for most of this gain at 33.49 points, while ProofWriter and CLRS improve by 2.85 points on average. Mechanism tests consistently favor anti-diamond supervision, whereas No-Orbit is statistically indistinguishable from full VCRT. The evidence does not establish a general benefit from exact orbit aggregation.