arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

感知、布局与验证:用于金融文档可靠直通处理的校准置信度

Perception, Layout, and Validation: Calibrated Confidence for Reliable Straight-Through Processing of Financial Documents

Yichao Jin, Yushuo Wang, Yuxuan Han, Kwan Ching Yee Sonia, Weiyang Song, Chiu Jin-Chun Kent, Wong Chong Hwee, Wong Tiong Kiat, Kenneth Zhu Ke, Jingyuan Zhao

arXiv 2609.20110首次发表:更新:

发表机构

OCBC(华侨银行)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种分解置信度层(感知、布局、验证)结合一致性风险控制,显著提升金融文档直通处理的校准可靠性,使自动批准率从0.1%-7.0%提升至49-72%。

AI 中文摘要

在无需人工审核的情况下,对从金融文档中提取的键值对进行直通处理(STP),需要校准的概率以及自动批准层级残差的有限保证。现代视觉语言模型(VLM)的出现提供了开箱即用的键值提取能力,但其口头化的置信度信号不可靠,且对字段正确性的追踪能力较弱。本文引入了一个分解置信度层,沿感知、布局和验证三个可解释通道展开。结合最终的一致性风险控制,该分数可用于金融文档的可靠直通处理。该方法在三个公共数据集上进行了验证,涵盖真实发票、合成发票和广告购买表单,使用了两个不同的VLM系列(Qwen3.6-27B和Gemini-3.1-Flash-Lite)。我们的分解分数持续改善了正确与错误提取的分离,将VLM口头化信号的AUROC从0.54-0.74大幅提升至0.90-0.99,且三个设计通道均有贡献。对工业部署至关重要的一点是,这实现了可用的直通处理。原生VLM置信度信号在目标错误率<10%的风险控制下仅能清除0.1%-7.0%的字段。相比之下,所提出的方法自动批准了49-72%的字段,同时将接受层的经验错误率保持在目标值或以下。

英文摘要

Straight-through processing (STP) on extracted key-value fields from financial documents without human review requires a calibrated probability together with a bounded guarantee on the residual error of the auto-approved tier. The emergence of modern Vision Language Models (VLMs) provides an out-of-the-box capability for extracting the key-values, but their verbalized confidence signals are unreliable and weakly track field correctness. This paper introduces a decomposed confidence layer along three interpretable channels, including perception, layout, and validation. Together with a final conformal risk control, the score can be used for reliable STP of financial documents. The method is validated on three public datasets covering real invoices, synthetic invoices, and ad-buy forms, using two different VLM families (Qwen3.6-27B and Gemini-3.1-Flash-Lite). Our decomposed score consistently improves the separation of correct from incorrect extractions, substantially raising the AUROC from 0.54-0.74 for VLM verbalized signals to 0.90-0.99 with contributions from all three designed channels. Crucially for industrial deployment, this enables usable STP. The native VLM confidence signals could clear only 0.1%-7.0% of fields under risk control at a target error of <10%. In contrast, the proposed method auto-approves 49-72% of fields while holding the empirical error of the accepted tier at or below the target.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑