发表机构
Institute of Computer Science, Warsaw University of Technology(华沙理工大学计算机科学研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对AI生成知识制品验证预算有限的问题,提出基于通过预测值(PVP)的贝叶斯有界验证模型,预测证据价值随规模增大而降低,并通过FEVEROUS和FEVER实验验证,提供维持目标PVP的最小预算计算。
AI 中文摘要
数字图书馆、知识库以及AI中介的知识服务日益依赖生成式系统来生成摘要、描述和其他多声明知识对象。然而,生成的速度远超验证能力:审阅者可能只需检查对象的一小部分声明,就需要决定该对象是否适合发表或下游使用。这造成了声明级验证与对象整体置信度之间的根本差距。因此,核心问题是:当对象规模增大而验证预算不变时,成功的部分检查对完整制品可靠性意味着什么。本文提出了一种以通过预测值(PVP)为中心的有界验证贝叶斯模型。该模型预测:在固定验证能力下,随着制品规模增大,证据价值会降低;增大验证预算会提升证据价值;当错误稀疏时,证据价值尤为脆弱。在FEVEROUS和FEVER上的受控实验支持这些预测,并表明在固定数量的虚假或未支持声明周围添加支持声明,可能使通过更可能发生,同时使通过的信息量减少。该模型进一步得出维持目标PVP所需的最小验证预算,为AI辅助的质量保证工作流程提供了一个实用组件,在此类流程中,生成的知识对象必须在发表或下游使用前接受检查。
英文摘要
Digital libraries, repositories, and AI-mediated knowledge services increasingly rely on generative systems to produce summaries, descriptions, and other multi-claim knowledge objects. Yet generation can scale far more easily than verification capacity: a reviewer may need to decide whether an object is suitable for publication or downstream use after checking only a small fraction of its claims. This creates a fundamental gap between claim-level verification and confidence in the object as a whole. The central question is therefore what a successful partial check implies about the reliability of the complete artifact when its size grows but the verification budget does not. This paper contributes a Bayesian model of bounded verification centered on the Predictive Value of Pass (PVP). The model predicts that evidentiary value decreases as artifacts grow under fixed verification capacity, improves with larger verification budgets, and is especially fragile when errors are sparse. Controlled experiments on FEVEROUS and FEVER support these predictions and show that adding supported claims around a fixed number of false or unsupported claims can make passing more likely while making a pass less informative. The model further yields the minimum verification budget required to maintain a target PVP, providing a practical component for AI-assisted quality-assurance workflows in which generated knowledge objects must be checked before publication or downstream use.
Comments11 pages, 2 figures. Accepted at the 28th International Conference on Asia-Pacific Digital Libraries (ICADL 2026)