arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

以更低成本拆分文档:面向LLM页面流分割的多分割边界决策

Splitting Documents at Lower Cost: Multi-Split Boundary Decisions for LLM-Based Page Stream Segmentation

Nikhil Reddy Pottanigari, Sepideh Kharaghani, Saverio Vadacchino, Alejandro Posada, Ying Zhang

arXiv 2609.22620首次发表:更新:

发表机构

ServiceNow Canada(ServiceNow加拿大公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出多分割边界决策(MSBD)方法,在单次调用中预测页面窗口内多个边界,以减少推理请求并提升零样本页面流分割的效率,实验表明在合适窗口下能兼顾准确性与效率。

AI 中文摘要

扫描邮件、上传的PDF和合并的附件通常以页面流的形式到达,必须在下游分类、提取或路由之前将其拆分为单独的文档。零样本大型语言模型可以在没有任务特定训练的情况下检测文档边界,但标准的页面分类(PC)和边界决策(BD)公式每次模型调用仅解析一个边界。我们引入了多分割边界决策(MSBD),它在单次调用中预测页面窗口内的多个边界,从而减少推理请求的数量。我们在多种语言模型、文档集合、输入模态和窗口大小上评估了MSBD。结果揭示了依赖于模型和语料库的工作范围,在该范围内MSBD保持了强大的分割准确性,同时显著提高了推理效率,随后在较大窗口时性能急剧下降。MSBD提供了最强的整体准确性-效率权衡,而大窗口暴露了不同模型间过分割和欠分割的明显行为。这些发现表明,当为目标语料库选择窗口大小时,多边界预测可以使零样本页面流分割更加高效。

英文摘要

Scanned mail, uploaded PDFs, and consolidated attachments often arrive as page streams that must be split into individual documents before downstream classification, extraction, or routing. Zero-shot large language models can detect document boundaries without task-specific training, but standard Page Classification (PC) and Boundary Decision (BD) formulations resolve only one boundary per model call. We introduce Multi-Split Boundary Decision (MSBD), which predicts multiple boundaries within a page window in a single call, reducing the number of inference requests. We evaluate MSBD across multiple language models, document collections, input modalities, and window sizes. The results reveal a model- and corpus-dependent operating range in which MSBD preserves strong segmentation accuracy while substantially improving inference efficiency, followed by a sharp decline at larger windows. MSBD provided the strongest overall accuracy--efficiency trade-off, while large windows expose distinct over- and under-segmentation behavior across models. These findings show that multi-boundary prediction can make zero-shot page stream segmentation more efficient when the window size is selected for the target corpus.

CommentsAccepted at the DocInsights Workshop at EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑