arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

评估评估者:面向大语言模型推理的验证自主等级(L0-L5)

Grading the Graders: Verification Autonomy Levels (L0-L5) for LLM Reasoning

Yajie Yin

arXiv 2608.19009首次发表:更新:

AI 中文总结

该研究提出验证自主等级(VAL)元标准,解决验证领域的等级概念混淆,明确不同验证方案的完备性边界,相关代码与评估材料已公开。

AI 中文摘要

大语言模型(LLM)越来越多地与验证器(步骤检查器、自一致性过滤器、基于工具的事实核查器、形式化证明助手)配对,这些验证器声称能够检测模型的错误。然而,验证文献中使用的“等级”一词至少有五种不同含义:验证粒度、概念抽象、风险层级、系统栈层以及真实值的认知来源。我们提出了验证自主等级(VAL)这一元标准,它沿单一轴对验证方案进行分类:验证规范来自何处,以及判定结果保证什么。VAL 的范围从 L0(LLM 自我声明,无确定性锚点)到 L2(客观真实值,仅保证正确性),再到 L3/L4(具有单属性或领域级完备性的可判定系统),而在无约束情况下 L5 不可能实现。VAL 的核心是完备性盲区:基于替换和采样的验证器可以确认所提出的候选成立,但无法证明未遗漏任何候选。我们进一步确定了文献中未明确提及的二分法:完备性仅对可形式化规范的属性可达,而经验性开放世界验证(事实核查、诊断)上限为锚定正确性(L2)。我们在四个领域(符号数学、行为监测、医学诊断和代码生成)以及现有最强的形式化验证基准中记录了这一点,其作者指出该验证器“关注每一步的正确性”。我们表明,粒度、概念层次结构、风险和系统栈的等级与 VAL 正交,解决了 17 篇被调查论文中存在的系统性混淆问题。代码和完整评估作为补充材料发布。

英文摘要

Large language models (LLMs) are increasingly paired with verifiers (step checkers, self-consistency filters, tool-based fact checkers, formal proof assistants) that claim to detect the model's errors. Yet the verification literature uses the word "level" to mean at least five different things: verification granularity, concept abstraction, risk tier, system-stack layer, and the epistemic source of the ground truth. We propose Verification Autonomy Levels (VAL), a meta-standard that classifies any verification scheme along a single axis: where does the verification spec come from, and what does the verdict guarantee? VAL ranges from L0 (LLM self-declaration; no deterministic anchor) through L2 (objective ground truth; correctness only) to L3/L4 (decidable systems with single-property or domain-level completeness), with L5 impossible in the unrestricted case. Central to VAL is the completeness blind spot: substitution- and sampling-based verifiers can confirm that proposed candidates hold, but cannot prove that no candidate was missed. We further identify a dichotomy the literature has not stated: completeness is reachable only for formally specifiable properties, whereas empirical open-world verification (fact-checking, diagnosis) caps at anchored correctness (L2). We document this gap empirically across four domains (symbolic mathematics, behavior monitoring, medical diagnosis, code generation) and in the strongest formal-verification baseline in our survey. We show that granularity, concept hierarchy, risk, and system stack are orthogonal to VAL, resolving a conflation across 17 surveyed papers. Code and the full literature assessment are released on Zenodo (DOI: 10.5281/zenodo.23120985).

Commentsv3: code and the full literature assessment released on Zenodo (DOI: 10.5281/zenodo.23120985); v2 added a reproducibility study (blind-LLM inter-rater kappa~0.8; human raters pending), an external-framework transfer test, and 15+ fixes. Writing was assisted by an AI language model; all experiments and research decisions are the author's own

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑