arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ViTeX-Bench:高保真视频场景文本编辑基准

ViTeX-Bench: Benchmarking High-Fidelity Video Scene Text Editing

Xinghao Chen, Xiangbo Gao, Jiongze Yu, Yuheng Wu, Zhengzhong Tu

arXiv 2609.40356首次发表:更新:

发表机构

Texas A&M University(德克萨斯A&M大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ViTeX-Bench提出视频场景文本编辑基准,含387个真实视频数据集和13指标三维评估协议,并发布ViTeX-Edit-14B编辑器,在文本正确性、时间稳定性和场景保留间取得最佳平衡。

AI 中文摘要

最近的视频生成越来越逼真且可控,但视频编辑仍不太成熟,尤其是在必须保留原始场景动态的精确局部编辑方面。视频场景文本编辑是指在场景表面(如店面招牌、白板和产品标签)替换文本,同时保留周围内容、运动和相机动态。尽管场景文本编辑在图像领域已得到充分研究,但实现高视觉质量、时间一致性和编辑局部性的视频场景文本编辑仍未得到充分探索。现有资源提供的配对真实视频数据有限,且通用视频编辑指标无法直接衡量所请求的文本是否随时间保持正确。我们引入了ViTeX-Bench,一个包含ViTeX-Dataset和三维评估协议的基准套件。该数据集包含387个真实世界的720p视频,带有文本区域掩码和编辑指令:其中230个提供经过审查的、由流程生成的配对编辑用于训练,157个构成冻结的评估分割。该协议通过13个指标评估文本正确性、视觉和时间质量以及编辑局部性,每个轴有一个主要指标,并对其权衡进行Pareto比较。OCR校准、人工评估和注释敏感性分析支持对这些分数的解读。在来自四个编辑家族的八个基线中,准确的文本、时间稳定性和场景保留仍然难以同时实现。我们还发布了ViTeX-Edit-14B,一个开源参考编辑器,在配对训练分割上进行了微调,并带有运动对齐的字形-视频条件。它在评估的视频原生编辑器中实现了CharAcc 0.688的最高平均值,并且在原始编辑器输出中实现了最低的可比较文本裁剪Warp。ViTeX-Bench为研究视频场景文本编辑中的这些权衡提供了一个可复现的基础。

英文摘要

Recent video generation is increasingly realistic and controllable, yet video editing remains less developed, particularly for precise local edits that must preserve the original scene dynamics. Video scene text editing replaces text on scene surfaces, such as storefront signs, whiteboards, and product labels, while preserving the surrounding content, motion, and camera dynamics. Although scene text editing is well studied for images, video scene text editing that achieves high visual quality, temporal consistency, and edit locality remains underexplored. Existing resources offer limited paired real-video data, and general video-editing metrics do not directly measure whether the requested text remains correct over time. We introduce ViTeX-Bench, a benchmark suite comprising ViTeX-Dataset and a three-axis evaluation protocol. The dataset contains 387 real-world 720p videos with text-region masks and editing instructions: 230 provide reviewed, pipeline-generated paired edits for training, and 157 form a frozen evaluation split. The protocol evaluates text correctness, visual and temporal quality, and edit locality through 13 metrics, with one primary metric per axis and a Pareto comparison of their trade-offs. OCR calibration, human evaluation, and annotation-sensitivity analyses support the interpretation of these scores. Across eight baselines from four editing families, accurate text, temporal stability, and scene preservation remain difficult to achieve together. We also release ViTeX-Edit-14B, an open-source reference editor fine-tuned on the paired training split with motion-aligned glyph-video conditioning. It achieves CharAcc 0.688, the highest mean among the evaluated video-native editors, and the lowest comparable text-crop Warp among raw editor outputs. ViTeX-Bench provides a reproducible foundation for studying these trade-offs in video scene text editing.

CommentsAccepted to NeurIPS 2026 (Evaluations and Datasets Track). 27 pages (10-page main text), 5 figures, 12 tables. Project page: https://vitex-bench.github.io/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑