arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.21792cs.AIcs.CVcs.IRcs.LG

HIRA:面向受监管行业文档分类的人在回路检索增强级联模型

HIRA: A Human-in-the-Loop Retrieval-Augmented Cascade for Document Classification in Regulated Industries

Shangxuan Tian, Yanhui Chen, Carlos Queiroz

首次发表
浏览论文内容

中文总结 AI 辅助

HIRA是面向受监管行业的无训练本地检索增强级联文档分类模型,结合多模态检索与本地LLM及人工审核,仅需少量人工修正即可大幅提升Macro-F1,减少LLM调用,性能接近全标注神谕模型。

中文摘要 AI 辅助

受监管行业的文档分类面临数据驻留限制、冷启动标签不足、审核能力稀缺以及模型治理流程成本高昂等问题。我们提出HIRA,这是一种面向受监管部署场景的无训练、本地部署的检索增强级联模型,通过验证校准的加权互反秩融合,将OCR文本上的BM25、稠密文本嵌入以及图像级表示相结合。置信度高的文档由检索直接分类;不确定或视觉上易混淆的文档会被传递给本地部署的LLM验证器,该验证器会接收OCR文本、检索到的示例、标签描述以及特定混淆术语。当验证器仍不确定时,文档会被提交给人工审核。每次修正会作为边际加权的检索示例存储,并更新狄利克雷平滑混淆图,使系统无需更新模型权重即可改进。在包含80个类别的私有贸易金融语料库上,HIRA处理全部30233份文档的生产流,仅请求对1945份文档(占比6.4%)进行人工修正,将Macro-F1从0.6218提升至0.8548。在修正后的Tobacco-3482基准上,使用本地部署的DeepSeek-R1-Distill-Qwen-32B验证器,HIRA达到0.9423的Macro-F1,比零样本LLM基线高出17.4个百分点,同时仅对约40%的文档调用验证器,减少了约60%的LLM调用量。通过518次人工修正(占池的24.8%),HIRA的性能与完全标注池的神谕模型相当,后者将全部2086份池文档用其真实标签索引。这些结果表明,在受监管部署场景中,选择性人工反馈与检索记忆自适应可作为长尾文档分类中反复模型重训练的实用替代方案。

英文摘要

Document classification in regulated industries is constrained by data residency, limited cold-start labels, scarce review capacity, and costly model-governance procedures. We present HIRA, a training-free, on-premises retrieval-augmented cascade for document classification in regulated deployments that combines BM25 over OCR text, dense text embeddings, and image-level representations through validation-calibrated weighted reciprocal-rank fusion. Confident documents are classified directly by retrieval; uncertain or visually confusable documents are passed to a locally hosted LLM verifier, which receives the OCR text, retrieved exemplars, label descriptions, and confusion-specific terms. When the verifier remains uncertain, the document is sent to human review. Each correction is stored as a margin-weighted retrieval exemplar and updates a Dirichlet-smoothed confusion graph, letting the system improve without updating model weights. On a private 80-class trade-finance corpus, HIRA processes the full 30,233-document production stream while requesting human correction for only 1,945 documents (6.4%), improving Macro-F1 from 0.6218 to 0.8548. On the corrected Tobacco-3482 benchmark, HIRA reaches 0.9423 Macro-F1 with a locally hosted DeepSeek-R1-Distill-Qwen-32B verifier, 17.4 percentage points above the zero-shot LLM baseline, while invoking the verifier for only about 40% of documents and reducing LLM calls by approximately 60%. With 518 human corrections (24.8% of the pool), HIRA matches the fully labelled pool oracle, in which all 2,086 pool documents are indexed with their ground-truth labels. These results show that selective human feedback and retrieval-memory adaptation can be a practical alternative to repeated model retraining for long-tail document classification in regulated deployments.

发表机构

  • Standard Chartered Bank(渣打银行)
  • OCBC(华侨银行)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑