arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.01499cs.CV

VTR-Bench:用于评估视频生成中视觉文本渲染的系统性基准

VTR-Bench: A Systematic Benchmark for Evaluating Visual Text Rendering in Video Generation

Yu Huang, Jungang Li, Zhiyuan Wang, Yonghua Hei, Song Dai, Jiayu Yang, Deyuan Liu, Xiang Zheng, Xiaoshuang Shi, Hao Cheng, Kaidi Xu

首次发表
浏览论文内容

中文总结 AI 辅助

VTR-Bench是一个系统性基准,通过300个提示和自动化评估流程,评估视频生成模型的视觉文本渲染能力,并引入关键帧引导的智能体框架,实验发现当前模型普遍存在文本渲染困难。

中文摘要 AI 辅助

最近的视频生成模型能够根据自然语言指令生成高度逼真的视频,其视觉质量接近电影级标准。然而,现有的评估基准主要评估视觉质量、美学吸引力和物理合理性,而对文本——这一在日常场景中传达信息的重要媒介——关注有限。生成的视频可能视觉上引人注目且具有逼真的主体,但场景内的文本渲染却可能出错。为解决这一被忽视的维度,我们引入了VTR-Bench,一个用于评估视频生成模型视觉文本渲染能力的系统性基准。VTR-Bench将文本置于具体应用场景中,如广告和科学视频,包含跨越五个场景类别的300个精心构建的提示。我们开发了一个与人类对齐的自动化评估流程,通过载体特定的转录分别评估文本保真度,并通过提示特定的查询链评估场景和运动要求。在评估之外,我们引入了一个关键帧引导的智能体框架,其中导演智能体协调图像和视频生成与视觉评估,通过视觉反馈指导迭代细化和候选选择。对11个最先进模型的实验揭示了在准确渲染场景文本方面的普遍困难,表现最佳的模型记录了0.250的总体词错误率(WER)。我们进一步分析了文本渲染失败,以描述当前视频生成模型面临的挑战。这些发现突显了视觉文本渲染是视频生成的关键挑战,并展示了一条实际的改进路径。代码可在该https URL获取。

英文摘要

Recent video generation models can produce highly realistic videos from natural language instructions, with visual quality approaching cinematic standards. Existing evaluation benchmarks, however, predominantly assess visual quality, aesthetic appeal and physical plausibility, while paying limited attention to text, an essential medium for conveying information in everyday scenes. A generated video may appear visually compelling and feature lifelike subjects, yet still render the text within the scene incorrectly. To address this overlooked dimension, we introduce \textbf{VTR-Bench}, a systematic benchmark for evaluating the \textbf{V}isual \textbf{T}ext \textbf{R}endering capabilities of video generation models. VTR-Bench situates text within concrete application scenarios, such as advertisements and scientific videos, with 300 carefully constructed prompts spanning five scenario categories. We develop an automated evaluation pipeline with human alignments that separately assesses text fidelity through carrier-specific transcription and scene and motion requirements through a prompt-specific chain of query. Beyond evaluation, we introduce a \textbf{Keyframe-Guided Agentic Framework} in which a Director agent coordinates image and video generation with visual evaluation, guiding iterative refinement and candidate selection through visual feedback. Experiments on 11 state-of-the-art models reveal widespread difficulties in accurately rendering scene text, with the best-performing model recording an overall word error rate (WER) of 0.250. We further analyze text rendering failures to characterize the challenges faced by current video generation models. These findings highlight visual text rendering as a key challenge for video generation and demonstrate a practical path toward improvement. Code is available at https://github.com/hardenyu21/VTR-Bench.

发表机构

  • City University of Hong Kong(香港城市大学)
  • The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
  • The Hong Kong Institute of AI for Science, City University of Hong Kong(香港城市大学香港人工智能科学研究院)
  • Westlake University(西湖大学)
  • University of Electronic Science and Technology of China(电子科技大学)
  • The Hong Kong University of Science and Technology(香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

↑