arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36598cs.CV

超越可读性:统一视频生成中视觉文本渲染与就地编辑的基准评测

Beyond Legibility: Benchmarking Visual Text Rendering and In-Place Editing in Unified Video Generation

Ziying Zhang, Litao Li, Junchao Liao, Tianyi Zeng, Siyu Zhu, Long Qin, Zhenghao Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

针对视频生成中视觉文本易崩溃且现有基准忽视时间动态的问题,提出统一诊断基准VidScribe,覆盖四种生成模式,评测11个系统,发现能力非单一且V2V编辑是瓶颈,并提供训练信号以优化文本生成。

中文摘要 AI 辅助

视频可以展现出令人信服的运动和照片级真实感,但当视觉文本崩溃时,它会立即失效。与通用场景内容不同,视觉文本在视频生成中是毫不宽容的:轻微的笔画损坏、时间不稳定或编辑错误会立刻破坏可读性和真实感。现有基准通过将文本视为偶然因素或使用忽略时间动态的静态OCR指标来忽视这一挑战。我们引入了VidScribe,一个统一的诊断基准,涵盖四种生成模式:从语言生成文本(T2V)、从参考图像转移文本身份(R2V)、在动态下维持文本(I2V)以及局部文本编辑(V2V)。VidScribe包含803个人工验证的样本,跨越一个12轴条件正交因子空间,涵盖内在文本属性、物理成像条件和时间行为。为了可靠评估,我们构建了一个基于轨迹的、门控的套件,包含11个共享指标和2个任务特定探针,在严格的可测量性条件下。对11个商业和开源系统的基准测试表明,视频文本能力是非单一性的,内容识别与笔画级字形正确性解耦。性能高度任务不对称:I2V最可靠地维持文本,而V2V编辑是主要瓶颈。出乎意料的是,退化集中在文本中心的结构和时间因素的一小部分,而不是不利的成像条件。进一步的探针显示,视觉参考提高了字形和排版保真度,而不是内容准确性,而局部编辑未能隔离目标文本,同时不破坏未声明的源文本。除了评估之外,VidScribe还提供了可操作的训练信号,其中与基准对齐的偏好优化可衡量地改善了视觉文本生成。此https URL。

英文摘要

A video can exhibit convincing motion and photorealism yet fail immediately when visual text collapses. Unlike generic scene content, visual text is unforgiving in video generation: minor stroke corruption, temporal instability, or editing errors instantly break legibility and realism. Existing benchmarks overlook this challenge by treating text as incidental or using static OCR metrics that ignore temporal dynamics. We introduce VidScribe, a unified diagnostic benchmark spanning four generation regimes: writing from language (T2V), transferring text identity from a reference (R2V), sustaining text under dynamics (I2V), and localized text editing (V2V). VidScribe contains 803 human-verified samples across a 12-axis conditionally orthogonal factor space covering Intrinsic Text Properties, Physical Imaging Conditions, and Temporal Behavior. For reliable evaluation, we build a track-grounded, gated suite with 11 shared metrics and 2 task-specific probes under strict measurability conditions. Benchmarking 11 commercial and open-source systems shows that video text capability is non-monolithic, with content recognition decoupled from stroke-level glyph correctness. Performance is highly task-asymmetric: I2V sustains text most reliably, whereas V2V editing is the primary bottleneck. Counter-intuitively, degradation concentrates on a small subset of text-centric structural and temporal factors rather than adverse imaging conditions. Further probes show that visual references improve glyph and typographic fidelity rather than content accuracy, while localized editing fails to isolate target text without corrupting undeclared source text. Beyond evaluation, VidScribe also provides an actionable training signal, where benchmark-aligned preference optimization measurably improves visual text generation. https://huggingface.co/datasets/Vicky0720/VidScribe.

发表机构

  • Alibaba Group(阿里巴巴集团)
  • Shanghai Jiao Tong University(上海交通大学)
  • Fudan University(复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

↑