arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ArmorOCR:基于观测迁移自蒸馏的 grounded 对抗性视觉感知

ArmorOCR: Grounded Adversarial Visual Perception via Observation-Transferred Self-Distillation

Linhan Cao, Siyuan Li, Jun Lan, Liangbo He, Guannan Li, Xiaolei Huang, Jun Jia, Shuheng Zhou, Huijia Zhu, Weiqiang Wang, Wei Sun

arXiv 2608.20122首次发表:更新:

AI 中文总结

本文提出ArmorOCR两阶段训练框架,结合OPSD与GRPO,推出含390幅图像的AdvSpot基准,可提升对抗性OCR感知能力并保留通用OCR性能

AI 中文摘要

大型多模态模型(LMM)已展现出强大的OCR识别能力,但仍易受对抗性视觉文本影响——这类文本对人类可读,却难以被模型定位与识别。现有OCR基准主要聚焦于自然或文档风格文本,而对抗性OCR评估在规模、任务覆盖或区域感知评估方面仍有限。本文将对抗性OCR定义为grounded OCR感知任务,并推出首个用于grounded对抗性OCR评估的基准AdvSpot。AdvSpot包含390幅带有区域级标注的图像,覆盖5个主要类别及13种细粒度对抗性OCR类型。为应对该挑战,本文提出ArmorOCR,这是一种用于鲁棒对抗性OCR感知的两阶段训练框架。ArmorOCR首先通过On-Policy自蒸馏(OPSD)从特权变换观测中获取缺失的对抗性OCR感知,随后通过带任务条件奖励的Group相对策略优化(GRPO)优化grounded OCR感知,任务条件奖励涵盖定位、识别、全检测及视觉问答(VQA)。在AdvSpot、其他对抗性OCR基准及通用OCR基准上的实验表明,ArmorOCR可持续提升对抗性OCR感知能力,同时保持具有竞争力的通用OCR性能。

英文摘要

Large multimodal models (LMMs) have demonstrated strong OCR recognition capabilities, yet remain vulnerable to adversarial visual text that is readable to humans but challenging for models to localize and recognize. Existing OCR benchmarks mainly focus on natural or document-style text, while adversarial OCR evaluations remain limited in scale, task coverage, or region-aware evaluation. In this paper, we formulate adversarial OCR as a \textbf{grounded OCR perception} task and introduce \textbf{AdvSpot}, the first benchmark for grounded adversarial OCR evaluation. AdvSpot comprises 390 images with region-level annotations, spanning 5 primary categories and 13 fine-grained adversarial OCR types. To address this challenge, we propose \textbf{ArmorOCR}, a two-stage training framework for robust adversarial OCR perception. ArmorOCR first acquires missing adversarial OCR perception from privileged transformed observations through On-Policy Self-Distillation (OPSD), and then refines grounded OCR perception through Group Relative Policy Optimization (GRPO) with task-conditioned rewards for localization, recognition, full spotting, and visual question answering (VQA). Experiments on our AdvSpot, other adversarial OCR benchmarks, and general OCR benchmarks demonstrate that ArmorOCR consistently improves adversarial OCR perception while preserving competitive general OCR capability.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑