PrismAlign:先验引导的多视角VLM对齐用于抗幻觉表格OCR
PrismAlign: Prior-Steered Multi-View VLM Alignment for Hallucination-Robust Table OCR
查看机构详情
- Huawei Technologies, Co., Ltd(华为技术有限公司)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
提出PrismAlign多VLM框架,利用表格逻辑先验和贝叶斯决策策略对齐多视角,减少表格OCR中的幻觉,在多个基准上达到最先进性能。
中文摘要 AI 辅助
表格提取常遭受频繁的结构错误和语义幻觉。我们提出PrismAlign,一个多VLM框架,通过对齐多样化的视觉视角来解决歧义。它整合表格逻辑的先验来评估输出合理性,将结构对齐与单元格内容对齐解耦。一种贝叶斯决策策略通过利用提取错误与可计算规则违反之间的相关性来最大化对齐准确性。在开源和自定义VLM上评估,PrismAlign减少了幻觉,并在OmniDocBench 1.5以及CC-OCR和PureDocBench的表格类别上达到了最先进的性能。
英文摘要
Table extraction suffers from frequent structural errors and semantic hallucinations. We propose PrismAlign, a multi-VLM framework aligning diverse visual perspectives to resolve ambiguity. It integrates priors of table logic to assess output plausibility, decoupling structural alignment from cell content alignment. A Bayesian decision strategy maximizes alignment accuracy by exploiting the correlation between extraction errors and computable rule violations. Evaluated on open-source and custom VLMs, PrismAlign reduces hallucinations and achieves state-of-the-art performance on OmniDocBench 1.5, as well as on the table category of CC-OCR and PureDocBench.