DocAnnot——利用生成式人工智能驱动的自动注释加速关键信息提取数据集的创建
DocAnnot -- Accelerating the Creation of Key Information Extraction Datasets with GenAI-Powered Auto-annotation
浏览论文内容
中文总结 AI 辅助
研究针对关键信息提取中训练数据集创建耗时的问题,提出DocAnnot框架,利用大型视觉语言模型、OCR及SICM算法,在基准测试中自动生成注释有一定F1分数,还能用于微调下游模型,大幅节省时间成本,减少人工依赖但未完全消除人工干预。
中文摘要 AI 辅助
关键信息提取(KIE)对许多文档应用至关重要,但传统上创建训练数据集是耗时的手动过程。我们引入了DocAnnot框架,它利用大型视觉语言模型进行标签值提取,利用光学字符识别进行文本/边界框检测,并采用一种新颖的空间感知上下文匹配(SICM)算法。SICM通过将空间关系和邻近分析与文本匹配相结合来改善标签值关联。我们在CORD和SROIE基准上评估该框架,其自动生成注释的F1分数分别为0.679和0.846。此外,我们研究了使用自动注释数据微调下游KIE模型的有效性。虽然人工注释数据仍然更优,但仅使用DocAnnot输出训练的模型也能达到可观的性能。结果表明,该框架虽大幅减少了对人工的依赖,但尚未完全消除人工干预。不过,它能使审阅者高效完善输出,实现接近完美的注释,比从头开始人工注释效率高得多,节省了大量时间和成本,对资源受限环境和快速模型原型制作很有价值。
英文摘要
Key Information Extraction (KIE) is vital for many document applications, but creating training datasets is traditionally a time-consuming manual process. We introduce DocAnnot, a framework that significantly accelerates KIE dataset generation. DocAnnot leverages a Large Vision Language Model (LVLM) for label value extraction, OCR for text/bounding box detection, and a novel Spatially Informed Contextual Matching (SICM) algorithm. SICM improves label-value association by combining spatial relationships and proximity analysis with textual matching. We evaluate our framework on the CORD and SROIE benchmarks, demonstrating its ability to auto-generate annotations with F1-scores of 0.679 and 0.846, respectively. Furthermore, we investigate the effectiveness of using auto-annotated data for fine-tuning downstream KIE models. While human-annotated data remains superior, models trained exclusively on DocAnnot's outputs attain respectable performance (e.g., LayoutLMv3 achieving an F1-score of 0.6765 on CORD). These results show that while our framework significantly reduces reliance on manual effort, it does not yet fully eliminate the need for human intervention. However, by automating the process to a point where reviewers can efficiently refine outputs, our system enables near-perfect annotations with much greater efficiency than manual annotation from scratch. This approach offers substantial time and cost savings, making it valuable for resource-constrained settings and rapid model prototyping.
发表机构
- Phi Labs, Quantiphi India(印度Quantiphi公司Phi实验室)
- Phi Labs, Quantiphi Canada(加拿大Quantiphi公司Phi实验室)
机构由 AI 辅助整理,请以论文原文为准。