DocPO: Advancing Document Policy Optimization via Tailored Step-Aware Rewards
DocPO:通过定制化的步骤感知奖励推进文档策略优化
机构 * Tencent Hunyuan(腾讯混元)
AI总结 本研究针对文档解析强化学习中奖励区分度不足的问题,提出 Step-Aware Annealing 机制,构建 DocPO 框架并在 OmniDocBench 等数据集上验证了其对 GRPO 式 RL 的性能提升效果。
Comments 14 pages. Accepted to the 34th ACM International Conference on Multimedia (ACM Multimedia 2026). Yunhao Wang and Binghong Wu contributed equally. Updated to the final camera-ready version with supplementary material