arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Theoria: 非正式推理状态上的重写-可接受性验证

Theoria: Rewrite-Acceptability Verification over Informal Reasoning States

Michael Saldivar, Ben Slivinski

arXiv 2607.01223首次发表:更新:

发表机构

Independent Researchers(独立研究者)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出Theoria验证架构,通过将候选解重写为带显式理由的序列化状态转换并验证变更完整性,在HLE-Verified Gold上以91.4%精确率认证105个问题,优于整体式LLM评判。

AI 中文摘要

何时应信任AI系统的答案?形式化证明助手提供确定性但无法覆盖大部分问题分布;标量LLM评判器提供覆盖率但产生不透明的分数,事后无法审计,且与任何LLM一样存在连贯性问题。我们提出Theoria,一种弥合这一差距的验证架构。候选解被重写为一系列带类型的状态转换,每个转换由显式理由(无论是引用、计算还是问题给定事实)授权,且每个转换均可独立审计。基础不变性是变更的完整性:连续证明状态之间的每个差异都必须被解释,因此隐藏前提作为未授权的突变浮现,而非静默通过。在HLE-Verified Gold(185个纯文本专家问题)上,Theoria以91.4%的严格精确率(Wilson 95% CI [84.5%, 95.4%])认证了105个问题。每个认证产生一个人类可读的证明轨迹,其中每一步都可被独立质疑。整体式LLM评判器在匹配覆盖率下达到可比的精确率,但在不同问题上失败(Jaccard 0.14-0.36),使得两种方法互补。在跨15个领域的95个对抗性中毒证明上,结构化评判器捕获94.7%,而整体式评判器为83.2%(p=0.0017)。整体11.5个百分点的差距集中在隐藏前提(90.6% vs. 62.5%,28个百分点差异)和伪造引用(100% vs. 90%)上,这些是形式分析预测优势的错误类别;在算术和定理误用错误上性能相同,这些错误上未预测到优势。在GPQA Diamond(n=65)上,认证精确率为97.1%(Wilson CI [85.1%, 99.5%])。

英文摘要

When should an AI system's answer be trusted? Formal proof assistants offer certainty but cannot reach most of the problem distribution; scalar LLM judges offer coverage but produce opaque scores that cannot be audited after the fact and are subject to the same coherence issues as any LLM. We present Theoria, a verification architecture that closes this gap. A candidate solution is rewritten into a sequence of typed state transitions, each licensed by an explicit justification, whether that be a citation, computation, or problem-given fact, and every transition is independently auditable. The foundational invariant is completeness of change: every difference between consecutive proof states must be accounted for, so hidden premises surface as unlicensed mutations rather than passing silently. On HLE-Verified Gold (185 text-only expert problems), Theoria certifies 105 at 91.4% strict precision (Wilson 95% CI [84.5%, 95.4%]). Every certification produces a human readable proof trace in which each step can be independently challenged. Holistic LLM judges achieve comparable precision at matched coverage but fail on different problems (Jaccard 0.14-0.36), making the approaches complementary. On 95 adversarial poisoned proofs across 15 domains, structured judges catch 94.7% versus 83.2% for holistic judging (p= 0.0017). The overall 11.5 pp gap concentrates in hidden premises (90.6% vs. 62.5%, a 28 pp difference) and fabricated citations (100% vs. 90%), the error classes where the formal analysis predicts an advantage; performance is identical on arithmetic and theorem-misapplication errors, where no advantage is predicted. On GPQA Diamond (n= 65), certified precision is 97.1% (Wilson CI [85.1%, 99.5%]).

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑