arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32477cs.CV

从清晰回溯:自我学习远距离文本识别

Back-Tracking from Clarity: Self-Learning to See Text from Afar

Duc-Tri Tran, Phi Le Nguyen, Minh Hoai

首次发表
浏览论文内容

中文总结 AI 辅助

提出自监督框架,通过视频回溯清晰文本检测生成伪标签,训练学生模型提升远距离小模糊文本的早期检测,并构建多语言数据集SceneText50验证效果。

中文摘要 AI 辅助

我们提出了一种自监督框架,旨在增强场景文本检测器在实例显示在显著距离(通常较小、模糊且常被传统模型遗漏)时识别和识别文本的能力。我们的方法利用现有文本点读模型在大而清晰的文本上的高保真性能作为基础监督者。通过视频序列时间回溯这些高置信度检测,我们自动为前帧合成伪标签,其中远距离文本仍视觉退化或尺寸不足。这些伪标签使得训练一个专门用于早期文本检测的学生模型成为可能,无需任何手动标注。该方法成功依赖于准确的伪标签生成,为此我们开发了一个专用场景文本追踪器,能够在具有挑战性的视频序列中保持一致的文本身份。此外,我们提出了SceneText50,一个多样化的多语言户外数据集,以促进训练和评估。实验表明,我们的框架显著提高了跨不同场景和语言的早期检测准确性和鲁棒性。代码和数据位于\href{this https URL}{this https URL}。

英文摘要

We propose a self-supervised framework designed to enhance the capability of scene text detectors in identifying and recognizing text in scenarios where instances are shown at significant distances, typically small, blurred, and frequently missed by conventional models. Our approach leverages the high-fidelity performance of existing text spotting models on large, clear text as a foundational supervisor. By temporally back-tracking these high-confidence detections through video sequences, we automatically synthesize pseudo-labels for preceding frames where the distant text is still visually degraded or undersized. These pseudo-labels enable training a student model specialized for early text detection, without requiring any manual annotation. The success of this approach depends on accurate pseudo-label generation, for which we develop a dedicated scene text tracker capable of maintaining consistent text identities across challenging video sequences. In addition, we propose SceneText50, a diverse multilingual outdoor dataset to facilitate training and evaluation. Experiments show that our framework significantly improves early detection accuracy and robustness across varied scenes and languages. Code and data are at \href{https://github.com/trid2912/BackTrackingText}{https://github.com/trid2912/BackTrackingText}.

↑