arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.00756cs.CL

联合训练还不够:面向多模态文档理解的条件跨粒度训练

Joint Training Is Not Enough: Conditioned Cross-Granularity Training for Multimodal Document Understanding

Chengguang Gan, Yunhao Liang, Hanjun Wei, Qinghao Zhang, Shiwen Ni

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对多模态文档理解任务,对比单任务、联合、条件训练三种方式,发现联合训练无强化效应,条件训练在两个语料库上实现跨粒度强化,还揭示了提示结构对性能的影响及条件训练的崩溃规避作用。

中文摘要 AI 辅助

互强化效应(MRE)探究当一个模型同时处理细粒度的跨度级任务和粗粒度的文档级任务时,二者是否能相互促进。我们在多模态文档理解任务上,使用三个语料库(两个收据语料库和一个扫描商业表单语料库)进行验证,对比了单任务训练、联合训练和条件训练三种方式;其中条件训练仅在训练时将某一粒度的黄金输出作为另一粒度的提示输入。我们构建了Doc-MRE,这是一个标注层,通过预注册的三名评审LLM委员会,将黄金字段提取(点级)与四个文档级方面(行级)配对,并通过盲重标注验证。预先设定的判断条件为:在相同的训练方案下,若某一机制在两个粒度上的表现均优于匹配的单任务模型,则称该机制具有强化作用。先前MRE研究假设的混合联合训练,在主要规模下未在任何语料库上产生强化效果:在CORD语料库上,其表现低于两个单任务模型;在另外两个语料库上,其表现为单任务式的“以某一粒度性能换取另一粒度性能”。条件训练在三个语料库中的两个上产生了强化效果:在CORD语料库上,点级性能提升0.5,行级性能提升4.8;在表单语料库上,点级性能提升7.2,行级性能提升11.0;在粗粒度侧的效果可明确验证,细粒度侧效果具有方向性,而在WildReceipt语料库上则出现性能交换。在该训练方案下,没有任何替代方法在任一粒度上的表现可显著优于条件训练。我们设置了两个字节相同提示的控制实验,以区分内容与格式的影响:打乱的条件输入会破坏粗粒度侧的性能,但对细粒度侧的性能影响小得多;中性内容控制实验在WildReceipt语料库上重现了全部细粒度增益,因此该增益源于提示结构,但在另外两个语料库上未发现可明确验证的效果。在表单语料库上,条件训练避免了模型崩溃:混合训练和中性内容控制实验均将50个测试文档全部分配给同一语义标签,而仅条件训练恢复了黄金分布。探针实验发现,所有训练机制下的信息均可解码,且条件训练下的信息解码能力未出现可明确验证的提升。

英文摘要

The Mutual Reinforcement Effect (MRE) asks whether a fine, span-level and a coarse, document-level task help each other when one model handles both. We test it in multimodal document understanding on three corpora, two of receipts and one of scanned business forms, comparing single-task, joint and conditioned training, which puts one granularity's gold output in the other's prompt during training only. We build Doc-MRE, an annotation layer pairing gold field extraction (point) with four document-level facets (line), from a three-judge LLM committee under a pre-registration, validated by blind re-annotation. One predicate, fixed in advance: at a shared recipe, a regime reinforces if it beats the matched single-task model on both granularities. Mixed joint training, the arrangement prior MRE work assumes, reinforces on no corpus at the main scale: it is below both single-task models on CORD and trades one granularity for the other on the two others, as single-task tuning does. Conditioned training reinforces on two of the three, CORD (+0.5 point, +4.8 line) and the forms corpus (+7.2 point, +11.0 line), resolvably on the coarse side and directionally on the fine one, and trades on WildReceipt; at that recipe no alternative measurably beats it on either side anywhere. Two byte-identical-prompt controls separate content from format: shuffled conditioning destroys the coarse-side skill but costs the fine side far less, and a neutral-content control reproduces the whole fine-side gain on WildReceipt, which is therefore prompt structure but buys nothing resolvable on the other two. On the forms corpus conditioning buys collapse avoidance: mixed training and the neutral control both assign the majority semantic label to all 50 test documents; only conditioning recovers the gold distribution. Probes find the information decodable under every regime with no resolvable increase under conditioning.

发表机构

  • University of Chinese Academy of Sciences(中国科学院大学)
  • Pusan National University(釜山国立大学)
  • Shenzhen University of Advanced Technology(深圳理工大学)

机构由 AI 辅助整理,请以论文原文为准。

↑