发表机构
T-Tech(T-Tech)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对受监管行业文档结构化字段提取的高成本问题,提出基于混合专家VLM的文档理解系统,通过难度感知数据整理与质量调整部署,实现成本大幅降低且性能领先可部署基线模型。
AI 中文摘要
在受监管行业中,每年从数亿份文档中提取结构化字段的成本仍然很高:定制化OCR级联仅覆盖部分工作流程,隐私规则禁止使用外部模型,而达到质量阈值的现有开源视觉语言模型(VLM)的部署成本高于人工标注。我们提出了一种已部署的文档理解系统,该系统基于混合专家VLM(总参数35B,激活参数3B)构建,在内部生产数据与通过难度感知流程整理的开放域文档混合数据上进行微调,该流程考虑了布局多样性、事实可提取性和跨模型一致性。该模型适配单块H100显卡,通过提示词服务于异构工作流程,在所有可部署的(非推理类)基线模型中领先,性能差距可达一个数量级。通过基于生产遥测数据校准的确认与修正成本开展质量调整成本分析显示,与人工基线相比,该模型可降低预期成本超过80%,与最佳竞争开源模型相比降低超过50%,而更大的基线模型在经济上仍不可行。
英文摘要
Extracting structured fields from hundreds of millions of documents annually remains costly in regulated industries: bespoke OCR cascades cover only a fraction of workflows, privacy rules preclude external models, and existing open-source VLMs that clear quality thresholds cost more to serve than human annotation. We present a deployed document-understanding system built on a Mixture-of-Experts VLM (35B total, 3B active), fine-tuned on in-house production data mixed with open-domain documents curated by a Difficulty-Aware pipeline for layout diversity, fact-extractability, and cross-model consistency. Fitting on a single H100 and serving heterogeneous workflows via prompting, the model leads all deployable (non-reasoning) baselines up to an order of magnitude larger. A quality-adjusted cost analysis, with confirmation and correction costs calibrated from production telemetry, shows it reduces expected costs by over 80% against the human baseline and by more than 50% against the best competing open-source model, while larger baselines remain economically unviable.