arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

你能信任置信度吗?面向文档提取的视觉语言模型ConfBench

Can You Trust the Confidence? ConfBench for Vision-Language Models on Document Extraction

Priyashree Roy, Sujitha Martin, Mohammad Rostami, Spencer Romo, Renhao Xue, Bob Strahan, Diego A. Socolinsky, Boyi Xie, Md Mofijul Islam

arXiv 2608.01792首次发表:更新:

发表机构

Amazon Web Services(亚马逊网络服务)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对视觉语言模型的文档提取任务推出校准专用基准ConfBench,评估多类VLM的置信度估计方法,揭示模态、模型能力等对置信度的影响,发布相关指标以助力可信IDP部署。

AI 中文摘要

智能文档处理(IDP)依赖视觉语言模型(VLM)提供足够可信的置信度分数,以在自动化处理和人工审核之间分配提取任务。现有文档基准以干净、高质量样本为主,低准确率区域样本过少,无法用于校准评估。我们推出ConfBench,这是首个针对关键信息提取(KIE)的校准专用基准,通过对多样化文档集应用20种受控退化流程构建,产生1346种变体和7万+实体级评估,覆盖全准确率范围。我们评估了4个专有和3个开放权重VLM,采用 verbalized(口头表述)和log-probability(对数概率)置信度估计方法,涵盖3种输入模态,发现:(i)OCR+图像模态能生成更准确的置信度估计;(ii)模型能力是主导因素:在Claude系列中,置信度质量随能力单调提升,而跨系列时参数数量是糟糕的预测指标;(iii)模型间校准质量差异大,从近乎完美到严重过度自信,模型专属事后校正会重新缩放绝对置信度值,用于基于阈值的路由,且不改变基于排名的操作指标;(iv)采用首词聚合的对数概率始终优于平均词和边际聚合。我们还推出ECARB,一种将判别增益转化为操作节省的审核预算指标。我们公开发布ConfBench,以支持对可信IDP应用部署的置信度估计器和校准方法的系统研究。

英文摘要

Intelligent document processing (IDP) with vision-language models (VLMs) hinges on confidence scores trustworthy enough to route extractions between automation and human review. Existing document benchmarks are dominated by clean, high-quality samples, leaving low accuracy regions too sparse for calibration assessment. We introduce ConfBench, the first calibration-specific benchmark for key information extraction (KIE), built by applying 20 controlled degradation pipelines to a diverse document set, yielding 1,346 variants and 70K+ entity-level evaluations spanning the full accuracy spectrum. We evaluate four proprietary and three open-weight VLMs under verbalized and log-probability confidence estimation methods across three input modalities, and find: (i) OCR+Image modality results in more accurate confidence estimates; (ii) model capability is the dominant factor: within the Claude family confidence quality scales monotonically with capability, while across families parameter count is a poor predictor; (iii) calibration quality varies widely across models, from near-perfect to severely overconfident, and per-model post-hoc correction rescales these absolute confidence values for threshold-based routing without altering ranking-based operational metrics; and (iv) log-probability with first-token aggregation consistently outperforms mean-token and margin aggregations. We also introduce ECARB, a review-budget metric translating discriminative gains into operational savings. We release ConfBench publicly to enable systematic study of confidence estimators and calibration methods for trustworthy IDP application deployment.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑