通过分类引导的大型视觉语言模型从文档中提取视觉信息
Visual Information Extraction from Documents via Classification-Guided Large Vision-Language Models
浏览论文内容
中文总结 AI 辅助
研究从视觉丰富文档提取视觉信息的挑战,提出分类引导大型视觉语言模型框架,通过解耦分类与提取、运用动态提示工程实现零样本推理,在真实数据集上性能超监督基线,为办公自动化文档理解提供高效方案。
中文摘要 AI 辅助
从视觉丰富的文档中提取视觉信息(VIE)具有挑战性。现有方法存在局限性。本文提出用于多类型VIE的分类引导大型视觉语言模型(LVLM)框架,将文档类型分类与内容提取解耦,采用基于上下文学习的动态提示工程注入特定任务知识,实现跨不同布局的零样本推理。理论上,该方法是一种条件计算形式。在有16种证书类型的真实世界投标数据集上评估,零样本方法性能优于强监督基线,特定领域微调可进一步提升性能,为办公自动化中的复杂文档理解提供高效可扩展解决方案。
英文摘要
Visual information extraction (VIE) from visually rich documents remains challenging due to high layout variability and real-world impairments. Existing methods typically rely on sequential OCR pipelines or end-to-end models requiring extensive labeled data and layout-specific training, limiting their scalability.We propose a classification-guided large vision-language model (LVLM) framework for multi-type VIE that achieves high accuracy with minimal supervision. The approach decouples document-type classification from content extraction and employs in-context learning (ICL)-based dynamic prompt engineering to inject task-specific knowledge, enabling robust zero-shot inference across diverse layouts. From a theoretical perspective, the proposed method can be viewed as a form of conditional computation that reduces task uncertainty and improves information efficiency during prompt-based inference. Evaluated on a real-world bidding dataset with 16 certificate types, our zero-shot method (based on Qwen2.5-VL-7B) outperforms a strong supervised baseline by 18.35 percentage points in F1-score (86.43\% vs. 68.08\%) and 0.23 in normalized edit distance (0.90 vs. 0.67). Optional domain-specific fine-tuning further improves performance to 93.65\% F1 and 0.93 NED, demonstrating superior robustness against seals, watermarks, and low contrast. The framework offers an efficient, scalable solution for complex document understanding in office automation. Code is available at https://github.com/FairmeHIT/Multi-VIE, and fine-tuned models at https://huggingface.co/fairme/Qwen2.5-VL-7B-SFT.
发表机构
- China Mobile Information Technology Co., Ltd.(中国移动信息技术有限公司)
- School of Information Science and Engineering, NingboTech University(宁波诺丁汉大学信息科学与工程学院)
机构由 AI 辅助整理,请以论文原文为准。