arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06532cs.CL

面向图表与文档理解的金融视觉语言模型的置信度估计

Confidence Estimation for Financial Vision-Language Models in Chart and Document Understanding

Reza Khanmohammadi, Simerjot Kaur, Charese H. Smiley, Ivan Brugere, Mohammad M. Ghassemi

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对金融视觉语言模型,评估7种置信度估计器的分布外迁移效果,发现仅训练得到的探测模型具备校准能力,其中接地感知探测模型可区分模型未使用图表生成的答案与流畅猜测,为金融场景的决策信任提供支撑。

中文摘要 AI 辅助

视觉语言模型(LVLMs)正越来越多地被用于读取金融图表、表格和文档,其中一个误读的数值可能会影响决策,而最具权威性的答案有时是模型未阅读相关材料就生成的。因此,实际的问题是信任而非准确性:哪些答案可以执行,哪些需要提交给审核人员。我们在5个开放权重LVLMs、3个金融视觉问答基准的4种条件(其中一个为双语)上评估了7种置信度估计器,包括3种仅推理的估计器和4种训练得到的内部探测模型;所有探测模型仅在自然图像上训练,未经过适配就应用于金融领域,因此结果衡量的是分布外迁移效果。得出三个核心发现:第一,稀缺的属性是校准而非排序:推理基线在将正确答案排在错误答案之上的竞争力尚可,但过度自信严重,校准误差远超阈值可容忍的范围,仅训练得到的探测模型能生成可用于阈值判断的分数;第二,可靠性是结构化的而非全局的,沿从业者可直接解读的两个轴变化:最佳估计器会随模型和任务变化,没有任何一个在20个(模型、条件)单元中领先超过8个,且受控的双语对比显示,看似的语言鲁棒性是组合产物,当单独解读模型时会消失;第三,在错误预算下被视为弃权(不执行)的情况下,可安全自动化的程度首先由模型的能力决定,仅由其置信度进一步收窄,因此弃权(不执行)能处理大部分最简单的条件,但在最困难的条件下几乎无法处理,在严格的5%预算下接近零。两种训练得到的探测模型具备弃权(不执行)策略所需的校准能力,其中仅接地感知探测模型会降低模型未使用图表就给出的答案的置信度,将检测到的非接地情况与流畅猜测区分开。

英文摘要

LVLMs are increasingly used to read financial charts, tables, and documents, where a single misread figure can move a decision and the most authoritative-looking answer is sometimes one the model produced without reading the exhibit. The operational question is therefore trust, not accuracy: which answers can be acted on, and which escalated to a reviewer. We evaluate seven confidence estimators, three inference-only and four trained internal probes, across five open-weight LVLMs and four conditions from three financial visual question-answering benchmarks, one bilingual; every probe is trained only on natural images and applied to finance without adaptation, so the results measure out-of-distribution transfer. Three findings hold. First, the scarce property is calibration, not ranking: the inference baselines rank correct above incorrect answers competitively but are badly overconfident, calibration error far above what a threshold can tolerate, and only the trained probes produce a thresholdable score. Second, reliability is structured rather than global, along two axes a practitioner can read directly: the best estimator shifts with both model and task, none leading more than eight of twenty (model, condition) cells, and a controlled bilingual contrast exposes an apparent language robustness as a composition artifact that dissolves once models are read one at a time. Third, cast as deferral under an error budget, how much can be safely automated is set first by the model's competence and only narrowed by its confidence, so deferral clears a real share of the easiest condition and almost none of the hardest, near zero at a strict 5% budget. Two trained probes carry the calibration a deferral policy needs, and among them only the grounding-aware one lowers its confidence on answers a model gives without using the figure, separating detected non-grounding from a fluent guess.

发表机构

  • Michigan State University(密歇根州立大学)
  • JPMorgan AI Research(摩根大通AI研究院)

机构由 AI 辅助整理,请以论文原文为准。

↑