HPD解析:分层并行文档解析
HPD-Parsing: Hierarchical Parallel Document Parsing
浏览论文内容
中文总结 AI 辅助
研究基于统一视觉语言模型的文档解析器顺序瓶颈问题,提出HPD解析,采用分层并行解码范式,含主要布局分支和P-MTP,实验显示其吞吐量高且精度有竞争力,为高效统一文档解析开辟新方向。
中文摘要 AI 辅助
高效协作通常将全局协调与并行执行相结合,这一原则在基于统一视觉语言模型(VLM)的文档解析器中尚未得到充分体现。现有统一解析器联合处理整个页面,但通过逐个令牌的自回归轨迹生成输出,形成随文档长度增长的顺序瓶颈。基于此,我们引入HPD解析,用分层并行解码范式取代全页面自回归生成。一个主要布局分支组织整体文档结构,动态将块级内容解码分配给并发分支,同时渐进多令牌预测(P-MTP)进一步减少每个分支内的解码步骤。在公共基准测试上的实验表明,HPD解析每秒实现4752个令牌,吞吐量是现有最快文档解析模型的2.62倍,是普通自回归基线的3.06倍,同时保持有竞争力的解析精度。这些结果确立了分层并行解码是全页面自回归生成的有效替代方案,为高效统一文档解析开辟了新方向。
英文摘要
Efficient teamwork typically combines global coordination with parallel execution, a principle not yet fully reflected in unified Vision-Language Model (VLM)-based document parsers. Existing unified parsers process an entire page jointly but generate its output through a single token-by-token autoregressive trajectory, creating a sequential bottleneck that grows with document length. Such full-page sequential generation overlooks a key property of document parsing: layout must be analyzed globally, whereas block content can be parsed in parallel. Based on this observation, we introduce HPD-Parsing, which replaces full-page autoregressive generation with a Hierarchical Parallel Decoding paradigm. A main layout branch organizes the overall document structure and dynamically assigns block-level content decoding to concurrent branches, while progressive multi-token prediction (P-MTP) further reduces the decoding steps within each branch. Experiments on public benchmarks show that HPD-Parsing achieves 4,752 tokens per second, delivering $2.62\times$ the throughput of the fastest existing document parsing model and $3.06\times$ that of the vanilla autoregressive baseline, while maintaining competitive parsing accuracy. These results establish hierarchical parallel decoding as an effective alternative to full-page autoregressive generation, opening a new direction for efficient unified document parsing.