arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

证明携带认知:通过现实结算奖励弥合验证差距

Proof-Carrying Cognition: Closing the Verification Gap with Reality-Settled Reward

Eshwar Reddy M, Sourav Karmakar

arXiv 2609.09776首次发表:更新:

发表机构

Testsigma University of San Diego; Intuit India(圣迭戈测试西格玛大学; 印度直觉公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出证明携带认知范式,通过现实结算奖励弥合验证差距,理论证明验证器质量与计算兑换率,实验显示健全验证器优于不健全者,并定义压力下健全性基准。

AI 中文摘要

语言模型推理的前沿进展源于对推理轨迹的强化学习,并集中在具有廉价且健全验证器的领域。我们认为该领域的根本约束是验证差距:在形式化领域之外,缺乏可扩展且不可腐蚀的奖励。我们做出四项贡献。(1) 理论:在best-of-N选择的联合高斯模型中,验证器-金标准相关性rho是测试时计算与能力之间的精确兑换率,不健全的验证器需付出多项式代价N^(1/rho^2);无边际的copula形式可预测真实LLM评判者的实际健全性,中位误差为4%。(2) 演示:在具有可执行真值的程序合成测试平台中,包括预注册的规模化复制,随着优化增强,不健全验证器在压力下健全性下降(在N=4096时从0.94降至0.32),而健全验证器单调提升;现实锚定结算在独立同分布和对抗压力下优于冻结验证器,将黑客差距从约0.27降至约0;健全性随结算标签对数线性扩展,在线策略结算比随机标记效率高约10倍。以真实LLM评判者和单元测试执行为金标准,弱评判者在best-of-N下失去健全性(p<0.001),强评判者更稳健,仅选择即可从诚实样本制造+0.53的黑客差距。在真实GRPO训练下,冻结奖励模型描绘完整过度优化曲线(执行奖励下降90%),而同一模型在10%结算流上重新拟合可保留6倍执行奖励。(3) 范式:证明携带认知,其中推理步骤是类型化概率声明,由仅在保留现实上训练的自建世界模型定价,并通过适当评分规则结算。(4) 基准:我们将压力下健全性指定为现实结算推理基准的首要指标。

英文摘要

Frontier gains in language-model reasoning come from reinforcement learning on reasoning traces and are concentrated in domains with a cheap, sound verifier. We argue the field's binding constraint is the verification gap: no scalable, incorruptible reward for reasoning outside formal domains. We make four contributions. (1) Theory: in a joint-Gaussian model of best-of-N selection, verifier-gold correlation rho is the exact exchange rate between test-time compute and capability, and an unsound verifier pays a polynomial penalty N^(1/rho^2); a margin-free copula form predicts realized soundness of real LLM judges to 4% median error. (2) Demonstration: in program-synthesis testbeds with executable ground truth, including a pre-registered scaled replication, unsound verifiers lose Soundness-under-Pressure as optimization grows (0.94 to 0.32 at N=4096) while a sound verifier improves monotonically; reality-anchored settlement beats a frozen verifier under i.i.d. and adversarial pressure, driving the hacking gap from ~0.27 to ~0; soundness scales log-linearly with settled labels, with on-policy settlement ~10x more label-efficient than random labeling. With real LLM judges and unit-test execution as gold, a weak judge loses soundness under best-of-N (p<0.001), a stronger judge is more robust, and selection alone manufactures +0.53 hacking gaps from honest samples. Under real GRPO training, a frozen reward model traces the full overoptimization curve (executed reward collapses 90%) while the same model refit on a 10% settlement stream preserves 6x the executed reward. (3) Paradigm: proof-carrying cognition, where reasoning steps are typed probabilistic claims priced by a self-built world model trained only on held-out reality and settled by proper scoring rules. (4) Benchmark: we specify Soundness-under-Pressure as the headline metric for a reality-settled reasoning benchmark.

Comments21 pages, 13 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑