arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过门控与衰减的在线策略蒸馏提升OCR忠实度

Improving OCR Faithfulness via Gated and Attenuated On-Policy Distillation

Baode Wang, Zuming Huang, Kexuan Ren, Jun Huang, Wei Chu

arXiv 2609.38282首次发表:更新:

AI 中文总结

针对视觉语言模型改写异常文本损害OCR忠实度的问题,提出GAD-RL,根据学生表现门控与衰减教师蒸馏,在Qwen3.5-2B上显著提升CHAOS-Bench与OmniDocBench性能。

AI 中文摘要

视觉语言模型可能会将图像中的异常文本改写为语言上合理的表达,从而损害OCR转录的忠实度。序列级任务奖励与局部教师指导是互补的,但随着学生模型的改进,来自同一教师的指导可能不再同样有效。离线分析表明,随着学生模型的改进,固定教师的监督变得越来越不利,这既体现在训练检查点之间,也体现在具有不同任务奖励的响应组之间。受此观察启发,我们引入了GAD-RL,它在联合后训练期间根据学生当前的任务表现和局部分布自适应地调节教师监督。冻结的教师以参考转录和学生生成的前缀为条件。GAD-RL对包含任务奖励至少为0.95的输出的响应组禁用蒸馏,并随着组平均奖励的增加而持续衰减蒸馏强度。它还将前向KL按学生对教师Top-1标记的概率进行加权,在学生对该候选的支持较低时调节局部辅助更新。在Qwen3.5-2B上,GAD-RL在CHAOS-Bench上实现了59.92%的Micro Recall,分别超过GRPO和GRPO+OPD(固定权重)8.45和4.43个百分点,同时在OmniDocBench v1.6上实现了91.18的Overall分数。

英文摘要

Vision-language models may rewrite anomalous text in images into linguistically plausible expressions, compromising OCR transcription faithfulness. Sequence-level task rewards and local teacher guidance are complementary, but guidance from the same teacher may not remain equally effective as the student improves. Offline analysis shows that supervision from a fixed teacher becomes progressively less favorable as the student improves, both across training checkpoints and across response groups with different task rewards. Motivated by this observation, we introduce GAD-RL, which adaptively regulates teacher supervision during joint post-training according to the student's current task performance and local distributions. A frozen teacher conditions on reference transcriptions and student-generated prefixes. GAD-RL disables distillation for response groups containing an output with task reward at least 0.95 and continuously attenuates distillation strength as group-mean reward increases. It also weights forward KL by the student's probability of the teacher's Top-1 token, moderating local auxiliary updates when student support for that candidate is low. On Qwen3.5-2B, GAD-RL achieves 59.92% Micro Recall on CHAOS-Bench, surpassing GRPO and GRPO+OPD (fixed-weight) by 8.45 and 4.43 percentage points, respectively, while achieving an Overall score of 91.18 on OmniDocBench v1.6.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑