发表机构
Stockmark Inc.(斯托克马克公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出Stockmark-Nemotron-3-Nano-Omni-JapanDocReader模型,通过能力注入与遗忘控制优化结构化文档解析性能,结合混合SFT与DAPO式RL取得优于SFT的效果。
AI 中文摘要
我们提出了Stockmark-Nemotron-3-Nano-Omni-JapanDocReader,这是一款基于Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16构建的日文文档理解模型。本研究的核心目标是通过能力注入与遗忘控制实现结构化文档解析:我们向一款面向推理的多模态模型中注入日文结构化文档解析能力,同时尽可能保留其文档VQA能力。我们研究了三种训练方式:仅使用结构化文档解析数据的解析中心式SFT、结合结构化文档解析与VQA数据的混合SFT、以及利用任务级奖励优化结构化解析的解析中心式RL。实验结果显示,解析中心式SFT可大幅提升结构化文档解析性能,但会造成可测量的VQA能力遗忘;混合SFT能缓解这种遗忘,同时保留几乎相同的结构化解析性能;在混合SFT检查点基础上应用基于DAPO的解析中心式RL,可进一步提升结构化解析性能,超越SFT的性能上限,最终得到发布的模型。训练数据由包含两个互补合成流的数据引擎构建而成:日文文档VQA流与程序化结构化文档解析流。我们还讨论了奖励设计与基于方差的提示过滤,以实现连续结构化文档解析奖励,强调了这些技术对让RL在长推理结构化文档解析任务中发挥效用的重要性。
英文摘要
We present Stockmark-Nemotron-3-Nano-Omni-JapanDocReader, a Japanese document understanding model built from Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16. The central goal of this work is structured document parsing via capability injection and forgetting control: we inject Japanese structured document parsing capability into a reasoning-oriented multimodal model while preserving its document VQA capability as much as possible. We study parsing-centric SFT, which uses only structured document parsing data; mixed SFT, which combines structured document parsing and VQA data; and parsing-centric RL, which optimizes structured parsing with a task-level reward. Our experiments show that parsing-centric SFT substantially improves structured document parsing performance but causes measurable VQA forgetting. Mixed SFT mitigates this forgetting while preserving nearly the same structured parsing performance. Applying DAPO-based parsing-centric RL on top of the mixed SFT checkpoint further improves structured document parsing beyond the SFT ceiling, producing the final released model. The training data is constructed with a data engine consisting of two complementary synthetic streams: a Japanese Document VQA Stream and a programmatic structured document parsing stream. We also discuss reward design and variance-based prompt filtering for continuous structured document parsing rewards, highlighting their importance for making RL effective in long-reasoning structured document parsing tasks.