arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ClearText-Video:连接视频修复与场景文本增强的大规模以文本为中心的视频数据集

ClearText-Video: A Large-Scale Text-Centric Video Dataset Bridging Video Restoration and Scene-Text Enhancement

Jinlong Li, Jiaming Ding, Dingfu Lu, Malcolm Hsiu, Chuang Ke, Kangning Yang, Bochen Guan, Lan Fu, Jie Cai, Huiming Sun, Zibo Meng

arXiv 2608.28784首次发表:更新:

发表机构

OPPO US AI Center; University of Wisconsin–Madison; University of California San Diego(OPPO美国人工智能中心; 威斯康星大学麦迪逊分校; 加利福尼亚大学圣迭戈分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究团队推出ClearText-Video数据集,评估发现视觉增强未必提升文本推理,仅OCR管道远逊于多模态推理,为相关系统提供基础。

AI 中文摘要

多模态大语言模型(MLLMs)近期在视觉-语言理解领域取得了显著进展,然而它们在以文本为中心的视频推理任务中的性能对输入质量高度敏感。现实世界中用户提供的视频常存在运动模糊、压缩伪影、噪声以及低分辨率文本等问题,这些问题会损害可靠的文本读取和下游推理能力。MLLMs能否在不同质量条件下稳健读取和推理现实场景文本仍是一个基础性开放问题。我们推出ClearText-Video(CTVid),这是一个大规模的、感知场景文本的基准,用于研究质量变化受控条件下以文本为中心的视频理解。CTVid包含4639个现实世界中富含文本的第一人称视角视频、55万余帧、160万个人工验证的场景文本标注,以及22万余个中英文空间/时序问答对。对于每个高质量视频,CTVid提供内容匹配的降质质量和恢复质量变体,支持两类任务:以文本为中心的视频修复和多质量视频问答(VideoQA)。我们在CTVid上评估了18种代表性修复方法和16种最先进的MLLMs。结果表明,视觉增强并不一定能保证文本保真度或下游推理增益:模糊比低分辨率的破坏性更强,恢复后的视频可能会改变MLLMs使用的文本证据,仅光学字符识别(OCR)的管道仍远落后于直接多模态推理。CTVid揭示了视频修复与文本基础理解之间的差距,为感知修复、质量稳健的以文本为中心的视频系统提供了严谨基础。

英文摘要

Multimodal Large Language Models (MLLMs) have recently made strong progress in visual--linguistic understanding. However, their performance on text-centric video reasoning remains highly sensitive to input quality. Real-world user-provided videos often contain motion blur, compression artifacts, noise, and low-resolution text, which impair reliable text reading and downstream reasoning. Whether MLLMs can robustly read and reason about real-world scene text under diverse quality conditions remains a fundamental open question. We introduce ClearText-Video (CTVid), a large-scale, scene-text-aware benchmark for studying text-centric video understanding under controlled quality variation. CTVid contains 4,639 real-world text-rich egocentric videos, 550K+ frames, 1.6M human-verified scene-text annotations, and 220K+ spatial/temporal question--answer pairs in Chinese and English. For each high-quality video, CTVid provides content-matched Degraded-Quality and Restored-Quality variants, supporting two task families: Text-Centric Video Restoration and Multi-Quality VideoQA. We evaluate 18 representative restoration methods and 16 state-of-the-art MLLMs on CTVid. The results show that visual enhancement does not guarantee textual fidelity or downstream reasoning gains: blur is more damaging than low resolution, restored videos can alter the textual evidence used by MLLMs, and OCR-only pipelines remain far below direct multimodal reasoning. CTVid exposes the gap between video restoration and text-grounded understanding, providing a rigorous foundation for restoration-aware, quality-robust text-centric video systems.

CommentsThis paper is accepted by 2026 Proceedings of the European Conference on Computer Vision

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑