arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

正确的信息抽取流水线取决于文档:小型本地模型的精度-能耗权衡

The Right Information Extraction Pipeline Depends on the Document: Accuracy-Energy Trade-offs for Small, Local Models

Christoph Walser, Mauricio Fadel Argerich, Jonathan Fürst

arXiv 2609.31341首次发表:更新:

AI 中文总结

本文研究小型本地模型在信息抽取中精度与能耗的权衡,发现批处理与FP8量化显著节能,且文档类型决定最优输入表示:版面丰富用视觉-语言模型,近纯文本用纯文本模型配廉价解析器。

AI 中文摘要

信息抽取流水线应处理页面图像还是解析文本,这取决于文档类型,答案在版面谱系的两端截然不同。我们在排除(封闭)云服务的约束下研究这一权衡:面向隐私敏感文档,在本地部署小型(≤80亿参数)纯文本模型和视觉-语言模型,并在涵盖输入表示、模型家族和推理配置的设计空间中,同时评估精度和能耗。在近纯文本的Kleister-NDA合同和版面丰富的VRDU表单上进行基准测试,我们发现批处理是主导的能耗杠杆,在不损失精度的情况下将每页能耗降低38-85%,而FP8量化在逐个请求服务时节省27-32%,但一旦应用批处理,每页节省不到1毫瓦时(9-19%)。预处理主导剩余能耗:神经OCR每页能耗是经典OCR的17倍,且从未达到帕累托前沿。哪种表示胜出随文档类型而翻转:视觉-语言模型在版面丰富的文档上胜出,而小型纯文本模型配合廉价解析器在近纯文本上胜出,后者在精度和成本上均优于任何视觉-语言配置。我们的工作为节能、合规的本地信息抽取提供了具体指南。

英文摘要

Whether an information extraction pipeline should process page images or parsed text depends on the document, and the answer flips across the layout spectrum. We study this trade-off under a constraint that rules out (closed) cloud services: privacy-sensitive documents processed on-premise by small ($\le 8\mathrm{B}$ parameter) text-only and vision--language models, evaluated on both accuracy and energy over a design space spanning input representation, model family, and inference configuration. Benchmarking on the near-plain-text Kleister-NDA contracts and the layout-rich VRDU forms, we find that batching is the dominant energy lever, cutting energy per page by 38-85% at no cost in accuracy, while FP8 quantization saves 27-32% when requests are served one at a time but less than 1mWh per page (9-19%) once batching is applied. Preprocessing dominates what remains: neural OCR costs $17\times$ more energy per page than classical OCR and never reaches the Pareto frontier. Which representation wins flips with the type of document: vision--language models on layout-rich documents and small text-only models with a cheap parser on near-plain text, where they are both more accurate and cheaper than any vision--language configuration. Our work yields concrete guidelines for energy-efficient, privacy-compliant local information extraction.

CommentsAccepted to DocInsights at EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑