arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.28421cs.AI

可验证奖励的程序学习:针对后训练大语言模型的符号反向传播

Program Learning with Verifiable Rewards: Symbolic Backpropagation for Post-Training LLMs

Vishvesh Bhat

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出PLVR方法,通过符号反向传播从输入输出示例学习显式程序,在代码基准上提升大语言模型推理能力,边际成本低且效果优于RL。

中文摘要 AI 辅助

对语言模型进行后训练以使其具备推理能力意味着更新其权重。监督微调与强化学习均将获得的能力置于模型内部,无法被检查、无法被逐步核验,也无法迁移至其他模型。本文提出,对于中间步骤可核验的任务,推理能力最好作为由确定性与神经原语组成的显式程序,置于基础模型权重之外。我们引入PLVR(Program Learning with Verifiable Rewards,可验证奖励的程序学习):一种直接从输入输出示例学习此类程序的后训练方法。其机制为符号反向传播:每个程序层携带类型化本体,在输出端针对真实值计算损失,并通过对原语签名进行类型推断反向传播所需输入本体,这是一种链式法则的类似物,其中信用分配是推导而非估计。RLVR核验最终结果,而PLVR的奖励是针对程序结构的逐步契约判决,具有密度。在LiveCodeBench v6与Tau2Bench上,采用PLVR的30B基础模型在匹配预算下平均比RL高出27.8个百分点,而规模大一个数量级的前沿模型则高出13.6个百分点。单个原语库可服务两个基准,因此新任务的边际成本为100个程序搜索示例,且无需新的微调数据。在相同预算下,将损失引导搜索替换为相同类型允许空间上的均匀采样,会将中位程序长度从65.6降至17.5,从而确定反向传递而非类型系统是该优势的来源。我们发布了符号反向传播库与一致性检查器,以便该方法可应用于除我们自身原语库之外的其他原语库。

英文摘要

Post training a language model to reason means updating its weights. Supervised finetuning and reinforcement learning both place the acquired capability inside the model where it cannot be inspected cannot be checked step by step and cannot be moved to another model. We argue that for tasks whose intermediate steps admit verification, reasoning is better placed outside the base models weights as an explicit program composed from deterministic and neural primitives. We introduce PLVR (Program Learning with Verifiable Rewards): a post training method that learns such programs directly from input-output examples. Its mechanism is symbolic backpropagation: each program layer carries a typed ontology a loss is computed at the output against ground truth and required input ontologies are propagated backward by type inference over primitive signatures: an analogue of the chain rule in which credit assignment is a derivation rather than an estimate. Where RLVR verifies a terminal outcome, PLVRs reward is a per step contract verdict dense over program structure. On LiveCodeBench v6 and Tau2Bench, 30B base models with PLVR outperform RL at matched budget by 27.8 points on average and frontier models an order of magnitude larger by 13.6 points. A single primitive library serves two benchmarks, so the marginal cost of a new task is 100 examples of program search and no new finetuning data. Replacing the loss guided search with uniform sampling over the same type admissible space at equal budget collapses the median program from 65.6 to 17.5, identifying the backward pass rather than the type system as the source of the advantage. We release the symbolic backpropagation library and a conformance checker so the method can be applied to primitive libraries other than our own.

发表机构

  • CoreThink AI

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑