arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06607cs.CLcs.IRcs.LG

用于高效成本的文档字段提取的推理前路由

Pre-Inference Routing for Cost-Efficient Document Field Extraction

Sreerekha Rajendran

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出推理前路由方法,利用文档级信号在低成本与强提取器间选择,可在保持提取质量的同时降低特定文档类型的提取成本,且仅在满足两条件时有效。

中文摘要 AI 辅助

大多数文档提取系统使用单一模型处理所有文档,这种方式简单但对简单文档而言成本过高,对复杂文档而言效果欠佳。本文研究是否能在提取前利用低成本的文档级信号预测文档的难度,以此在低成本提取器和强提取器之间做选择。研究发现,仅当两个条件同时满足时,路由才会有效:一是低成本模型的失败频率足够高,使路由具备价值;二是这些失败可通过图像质量、布局等可见特征预测。将上述条件转化为实用测试后,应用于五种文档类型:当两个条件均满足时,校准后的路由器能将收据的成本降低31%-33%,退化的广告采购表单的成本降低77%,同时保持的F1值与始终选择大型模型的情况仅相差0.02;若任一条件不满足,路由则无效,例如清晰的数字发票或营养标签这类本身易读取的文档。小型标记试点可预测路由是否有效,在首次测试的两种场景中,预测均正确。简单的词袋路由器效果与工程特征相当,表明主要限制因素是文档类型而非路由器设计;研究使用可解释特征解释哪些类型的文档可路由。路由器需针对每个数据集重新训练,无法跨数据集迁移,即使是同一类型的数据集也不例外。上述结果在成本差异为5倍和3倍的两组模型对中均成立。

英文摘要

Most document-extraction systems use a single model for all documents. This is simple but can be costly for easy cases and less effective for difficult ones. We examine whether we can predict a document's difficulty before extraction using inexpensive, document-based signals, and use this to choose between a cheaper and a stronger extractor. We find that routing only helps if two conditions hold: the cheaper model fails often enough to make routing worthwhile, and those failures can be predicted from visible features such as image quality and layout. We turn these into a practical test and apply it to five genres. When both conditions are met, the calibrated router reduces cost by 31-33% on receipts and 77% on degraded ad-buy forms while keeping quality within 0.02 F1 of always choosing the large model. Routing does not help if either condition is missing, as with clean digital invoices or nutrition labels that are already easy to read. A small labeled pilot can predict whether routing will work, and in the two cases where we ran it first, the prediction was correct. A simple bag-of-words router works about as well as engineered features, showing that the main limit is the genre, not the router design; we use interpretable features to help explain which genres can be routed. The router must be retrained for each dataset and does not transfer across datasets, even within the same genre. These results hold for two model pairs with cost differences of 5x and 3x.

补充信息

↑