AI 中文总结
研究针对文档分类中资源利用低效问题,提出分层强化学习框架DocHRL,将其作为顺序决策问题,以负总预期成本为奖励信号,经近端策略优化训练,在提升分类性能同时降低文档处理成本。
AI 中文摘要
现实世界的文档分类管道通常对每个传入文档应用相同的模型序列,而不考虑其复杂性或类型。这导致计算和人力资源利用效率低下。我们引入了DocHRL,一个分层强化学习框架,它学习在每个文档基础上自适应地动态选择最具成本效益的分类策略。DocHRL将文档分类表述为具有两级策略层次结构的顺序决策问题。奖励信号是负的总预期成本。在RVL-CDIP基准上用近端策略优化训练后,DocHRL在16个文档类别上实现了0.973的宏F1,同时将平均每个文档成本降低到2.74个归一化单位。结果表明,成本感知强化学习可同时提高文档理解系统的分类性能和运营效率。
英文摘要
Real-world document classification pipelines typically apply the same sequence of models to every incoming document, regardless of its complexity or type. This leads to inefficient use of compute and human resources: simple documents are over-processed while difficult ones may not receive enough scrutiny. We introduce DocHRL, a hierarchical reinforcement learning framework that learns to adaptively and dynamically select the most cost-effective classification policy on a per-document basis. DocHRL formulates document classification as a sequential decision problem with a two-level policy hierarchy: a top-level policy selects among broad options (vision classifiers, LLMs, OCR, and human-in-the-loop review), while option-specific sub-policies choose the concrete model or tool to invoke. The reward signal is the negative total expected cost, which captures inference cost, cost of misclassification, and cost of human labelling. Trained with Proximal Policy Optimisation on the RVL-CDIP benchmark, DocHRL achieves a macro F1 of 0.973 across 16 document classes while reducing average per-document cost to 2.74 normalised units compared to substantially higher costs incurred by fixed standalone classifiers. Our results demonstrate that cost-aware reinforcement learning can simultaneously improve classification performance and operational efficiency in document understanding systems.
Comments9 pages, 4 tables, 2 figures