发表机构
Tencent Hunyuan(腾讯混元)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对文档解析强化学习中奖励区分度不足的问题,提出 Step-Aware Annealing 机制,构建 DocPO 框架并在 OmniDocBench 等数据集上验证了其对 GRPO 式 RL 的性能提升效果。
AI 中文摘要
用于文档解析的强化学习(RL)通常依赖于基于编辑距离的参考奖励(例如树编辑距离),但在高精度 regime 下优化仍存在困难,因为此类奖励的区分度较弱:接近正确的输出会获得非常相似的分数,难以给困难样本提供充足的学习信号。我们提出 Step-Aware Annealing(SAA),一种即插即用的奖励锐化机制,在训练过程中逐步提升奖励曲率,放大高分样本间的细微质量差异,同时在早期学习阶段保持稳定性。基于 SAA,我们引入 DocPO,这是一种文档策略优化框架,其元素特定的参考奖励锚定编辑距离信号:文本使用归一化字符串编辑距离(NED),表格使用树编辑距离相似度(TEDS),公式使用混合 Rubric+edit 奖励。在 OmniDocBench 和 DocElemHard 上的实验表明,SAA 能持续提升 GRPO 式 RL 在各类文档元素上的性能,且无需为奖励构建额外的人工监督。
英文摘要
Reinforcement learning (RL) for document parsing often relies on reference-based rewards rooted in edit distance (e.g., tree edit distance), yet it remains hard to optimize in the high-accuracy regime because such rewards become weakly discriminative: near-correct outputs receive very similar scores, providing limited learning signal for hard cases. We propose Step-Aware Annealing (SAA), a plug-and-play reward sharpening mechanism that progressively increases reward curvature during training, amplifying subtle quality differences among high-scoring samples while preserving stability in early learning. Built on SAA, we introduce DocPO, a document policy optimization framework with element-specific, reference-based rewards anchored by edit-distance signals: normalized string edit distance (NED) for text, tree edit distance similarity (TEDS) for tables, and a hybrid Rubric+edit reward for formulas. Experiments on OmniDocBench and DocElemHard show that SAA consistently improves GRPO-style RL across document elements over non-annealed rewards, without requiring additional human supervision for reward construction.
Comments14 pages. Accepted to the 34th ACM International Conference on Multimedia (ACM Multimedia 2026). Yunhao Wang and Binghong Wu contributed equally. Updated to the final camera-ready version with supplementary material