发表机构
Hitachi, Ltd.; Hitachi Rail(日立制作所; 日立铁路)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究指出AI评估仅靠正确性不足,提出需纳入验证成本,定义了验证成本错误,通过代码生成等证据说明高基准准确率可能掩盖高验证成本,主张评估应考虑现实资源约束下的错误可检测性。
AI 中文摘要
AI生成模型的可靠性通常通过输出正确性来衡量,但在实际应用中,其可靠性取决于验证这些输出所需的精力。我们认为当前的评估指标忽略了一种关键失效模式:验证成本错误(Verification-Cost Errors,VCEs),其定义为在特定部署场景的验证预算内,一定比例的验证者无法识别的错误输入输出对。与标准的“幻觉”概念不同,VCEs是操作性定义的,由在预算内无法正确识别而非输出本身的任何属性决定。我们假设可信度和权威呈现是导致这种失效的因素,而非定义条件。为捕捉这种不对称性,我们引入相对于部署预算的验证成本概念,作为当前评估未常规捕捉的操作性维度,该概念是一种概念工具而非最终指标。来自代码生成和多模态文档理解的证据表明,高基准准确率在实践中可能掩盖了大量的验证精力。因此,我们认为仅用正确性作为可靠性衡量标准是不够的,AI评估应明确考虑验证成本,以反映在现实资源约束下错误是否能被检测到。
英文摘要
The reliability of AI generative models is typically measured by output correctness, yet in practice it depends on the effort required to verify those outputs. We argue that current evaluation metrics overlook a critical failure mode: Verification-Cost Errors (VCEs), defined as incorrect input-output pairs that a declared fraction of the verifier population fails to identify within the verification budget available in a given deployment context. Unlike standard notions of "hallucination", VCEs are defined operationally, by the failure of correct identification within budget rather than by any property of the output itself. Plausibility and authoritative presentation are hypothesised contributors to that failure, not defining conditions. To capture this asymmetry, we introduce the notion of verification cost relative to a deployment budget as an operational dimension that current evaluation does not routinely capture. The quantity is presented as a conceptual instrument rather than a finalized metric. Evidence from code generation and multi-modal document understanding shows that high benchmark accuracy can mask significant verification effort in practice. We therefore take the position that correctness alone is insufficient as a measure of reliability. AI evaluation should explicitly account for verification cost, reflecting whether errors can be detected under realistic resource constraints.
Comments20 pages, 1 table