发表机构
IBM Consulting(IBM咨询)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出 FRAGMENT 框架,将文档表示为类型化关系图并因式分解分布为结构与内容的乘积,经两阶段生成可用于文档生成、编辑及伪造检测,在多个数据集上验证了其性能。
AI 中文摘要
发票、表单、报告、科学文章等结构化文档的意义源于空间布局、文本内容与逻辑结构的相互作用。像素级或 token 级的生成模型往往难以有效捕捉这些依赖关系。我们探索 FRAGMENT,这一将文档表示为类型化关系图的生成框架,其将分布因式分解为 p(结构, 内容) = p(结构) * p(内容 | 结构)。该框架包含两个阶段:第一阶段是 Architect,它是一个受文档类别条件约束的因果掩码 Transformer,以自回归方式生成图拓扑结构与类型化空间关系;第二阶段是 Builder,它是一个基于 GATv2 的图注意力网络,为图补充归一化边界框、文本及视觉风格属性。两个阶段均定义了显式似然模型,产生可处理的文档级似然,可作为伪造检测的异常分数。对于可控编辑,基于提示的扩展通过交叉注意力将指令嵌入注入 Builder,实现语义与实体感知的修改。我们描述了在 DocLayNet 上的训练,以及在 FUNSD 和 SROIE 上的微调。在 DocLayNet、FUNSD 和 SROIE 上的实验,将 FRAGMENT 与代表性的自回归、仅布局及基于图的基线进行了对比,提供了所提出的因子化解码图生成框架的特性与权衡的实证分析。
英文摘要
Structured documents such as invoices, forms, reports, and scientific articles derive meaning from the interplay between spatial layout, textual content, and logical structure. Generative models operating at the pixel or token level often struggle to capture these dependencies effectively. We explore FRAGMENT, a generative framework that represents a document as a typed relational graph and factorizes its distribution as p(structure, content) = p(structure) * p(content | structure). The framework consists of two stages. The first stage, the Architect, is a causally masked Transformer conditioned on document category that autoregressively generates the graph topology and typed spatial relations. The second stage, the Builder, is a GATv2-based graph attention network that enriches the graph with normalized bounding boxes, text, and visual style attributes. Both stages define explicit likelihood models, yielding a tractable document-level likelihood that serves as an anomaly score for forgery detection. For controlled editing, a prompt-conditioned extension injects instruction embeddings into the Builder through cross-attention, enabling semantic and entity-aware modifications. We describe training on DocLayNet and fine-tuning on FUNSD and SROIE. Experiments on DocLayNet, FUNSD, and SROIE evaluate FRAGMENT alongside representative autoregressive, layout-only, and graph-based baselines, providing an empirical analysis of the characteristics and trade-offs of the proposed factorized graph generation framework.